Power distribution equipment sub-pixel precision positioning method, system, device and storage medium
Patent Information
- Application Number
- CN202611248383.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-18
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]上述方案虽然在工况稳定的场景下能够满足基本需求,但仍存在多处不足,一方面是固定类型特征算子鲁棒性不足,SIFT等固定特征算子针对不同类型特征(角点、圆形按键、直线边缘、阵列中心)的适应性有限,当设备因磨损、污染或遮挡导致特征外观变化时,SIFT描述子匹配错误率显著上升,导致亚像素精化阶段坐标偏差突增;另一方面,现有SLAM辅助方案以目标区域点云的深度均值与标准差约束EPnP/PnP解的深度分量(一维标量约束),无法修正旋转角估算偏差
[0015]本发明提供了一种配电设备亚像素精密定位方法、系统、设备和存储介质,通过条件化神经特征场取代固定特征算子,以数据驱动方式统一解决各类控制点的亚像素定位问题,能够消除特征类型误判;通过基于深度法线联合先验约束PnP解的平移深度分量与旋转法线方向,能够消除深度旋转耦合歧义,提高位姿估计的精准度;通过卡尔曼滤波预测前馈运动补偿,提升二维坐标的准确性;通过增量式外观自适应更新,使系统长期保持高精度;本发明能够实现面向配电设备精密操控的高精度视觉定位。
Smart Images

Figure CN122820845A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of equipment positioning technology, and in particular to a sub-pixel precision positioning method, system, device, and storage medium for power distribution equipment. Background Technology
[0002] Currently, "leg-arm" composite inspection platforms, using quadruped robots as carriers and equipped with multi-degree-of-freedom robotic arms, have become an important technological direction for automated inspection of power distribution networks. The absolute positioning accuracy of the robotic arm's end effector must reach the sub-millimeter level. Existing mainstream solutions use the PnP algorithm combined with SIFT feature point matching to achieve three-dimensional spatial positioning of key targets on the equipment, further improve positioning accuracy with sub-pixel corner detection, and select the EPnP algorithm to improve the speed of attitude estimation.
[0003] While the aforementioned solutions can meet basic requirements in stable operating scenarios, they still have several shortcomings. Firstly, the robustness of fixed-type feature operators is insufficient. Fixed feature operators such as SIFT have limited adaptability to different feature types (corners, circular buttons, straight edges, array centers). When the appearance of features changes due to wear, contamination, or occlusion, the SIFT descriptor matching error rate increases significantly, leading to a sudden increase in coordinate deviation during the sub-pixel refinement stage. Secondly, existing SLAM-assisted solutions constrain the depth component of the EPnP / PnP solution (a one-dimensional scalar constraint) using the mean and standard deviation of the target region's point cloud depth, failing to correct rotation angle estimation bias. In frontal proximity scenarios, depth estimation bias and rotation angle bias are strongly coupled; large rotation errors cause the end effector position deviation to exceed the acceptable range, and the EPnP one-dimensional depth constraint cannot eliminate the ambiguity of depth-rotation coupling. Furthermore, existing solutions weight the acquisition frames through stability scoring, failing to compensate for residual micro-vibrations at the moment of acquisition. When uneven ground causes periodic vibrations, the waiting time for the stabilization window is long and irregular, resulting in low effective acquisition efficiency. Summary of the Invention
[0004] To address the aforementioned technical problems, this invention provides a subpixel precision positioning method, system, device, and storage medium for power distribution equipment, which can achieve sub-millimeter-level three-dimensional positioning, thereby meeting the actual needs for high-precision positioning in automated operation and maintenance of power distribution networks.
[0005] In a first aspect, the present invention provides a sub-pixel precision positioning method for power distribution equipment, the method comprising: The frame images of the area to be localized, collected by a quadruped robot, are identified to obtain the type of power distribution equipment and the image bounding box. The corresponding three-dimensional control point model and type embedding vector are retrieved from the database. The encoded frame image coordinates and the type embedding vector are concatenated and input into the conditional neural feature field to obtain scalar saliency values. The scalar saliency values within the image bounding box are then densely sampled to obtain a saliency density map. The conditional neural feature field is constructed based on a fully connected network. Using the local maxima of the saliency density map as candidate control points, and performing gradient ascent iteration on the two-dimensional coordinates of each candidate control point, sub-pixel precision two-dimensional coordinates of each candidate control point are obtained. Obtain the truncated signed distance field of the region to be localized, and extract the depth value and gradient direction of the zero-level voxel set within the voxel region corresponding to the image bounding box to construct a joint depth normal prior; Based on the three-dimensional control point model, the sub-pixel precision two-dimensional coordinates, and the joint prior of the depth normal, a weighted constrained PnP objective function is constructed, and the estimated pose of the power distribution equipment is obtained by solving it.
[0006] Further, the step of concatenating the encoded frame image coordinates and the type embedding vector and inputting them into the conditional neural feature field to obtain scalar saliency values, and then densely sampling the scalar saliency values within the image bounding box to obtain a saliency density map includes: Fourier position encoding is performed on the frame image coordinates, and the encoded frame image coordinates and the type embedding vector are concatenated and input into the conditional neural feature field to obtain scalar significance values; The coordinates within the image bounding box are densely sampled with a preset step size, and the scalar saliency values are queried in batches to obtain a saliency density map; The conditional neural feature field is constructed based on a multi-layer fully connected network, using a sine function as the activation function, and employing mean squared error loss based on focus weights and information noise contrast loss as loss functions.
[0007] Further, the steps of obtaining the truncated signed distance field of the region to be localized, and extracting the depth value and gradient direction of the zero-level voxel set within the voxel region corresponding to the image bounding box, and constructing a joint depth-normal prior include: The truncated signed distance field of the region to be localized is obtained by SLAM technology, and the zero-level voxel set is extracted within the voxel region corresponding to the image bounding box. Extract the depth value set of each voxel in the zero-level voxel set in the camera coordinate system, and calculate the mean and standard deviation of the depth value set to obtain the depth prior mean and depth prior standard deviation. Extract the set of surface normal vectors from the zero-level collective element set, calculate the mean direction of the normal using the averaging method based on geodesic distance, and calculate the angular standard deviation of the angle between each normal and the mean direction of the normal to obtain the a prior mean direction of the normal and the a prior angular standard deviation of the normal.
[0008] Furthermore, the step of constructing a weighted constrained PnP objective function based on the three-dimensional control point model, the sub-pixel precision two-dimensional coordinates, and the joint prior of the depth normal includes: Based on the three-dimensional control point model and the sub-pixel precision two-dimensional coordinates, a reprojection error term is constructed, wherein the coordinate point weights in the reprojection error term adopt the scalar significance values of the coordinate point positions. Based on the depth value of the translation vector and the prior mean of the depth, a depth regularization term is constructed, and based on the direction of the frontal normal vector of the device model coordinate system and the prior mean of the normal, a normal regularization term is constructed. The objective function of the weighted constraint PnP is to minimize the weighted sum of the reprojection error term, the depth regularization term, and the normal regularization term. The weight of the depth regularization term is determined based on the prior standard deviation of the depth, and the weight of the normal regularization term is determined based on the prior standard deviation of the normal angle.
[0009] Furthermore, prior to the step of recognizing the frame images of the region to be localized acquired by the quadruped robot, the method further includes: Based on the triaxial force data and horizontal acceleration data of the quadruped robot's feet, a dynamic model of the quadruped robot based on body vibration was determined through data fitting. Based on the dynamic model, Kalman filtering is used to predict the planar displacement of the frame image to obtain the planar micro-displacement vector; The vibration amplitude is calculated based on the planar micro-displacement vector, and the vibration amplitude is compared with an amplitude threshold. In response to the vibration amplitude being less than or equal to the amplitude threshold, the quadruped robot is triggered to acquire frame images of the area to be located.
[0010] Furthermore, after the step of obtaining the sub-pixel precision two-dimensional coordinates of each candidate control point, the method further includes: The sub-pixel precision two-dimensional coordinates of each candidate control point are feedforward corrected based on the planar micro-displacement vector to obtain the corrected two-dimensional coordinates.
[0011] Furthermore, after the step of calculating the estimated pose of the power distribution equipment, the method further includes: Based on the estimated pose, calculate the reprojection residual of each 3D control point, and count the proportion of 3D control points whose reprojection residual exceeds the residual threshold. Based on the comparison between the ratio and the appearance drift threshold, it is determined whether appearance drift exists. The appearance drift threshold is retrieved from the database. If so, the parameters of the last layer of the conditional neural feature field are updated incrementally with microsteps driven by the reprojection residual gradient, and the 3D control points are rolled back and verified with geometric distance error as a hard constraint.
[0012] In a second aspect, the present invention provides a sub-pixel precision positioning system for power distribution equipment, the system comprising: The control point extraction module is used to identify the frame images of the area to be localized based on the quadruped robot, obtain the type of power distribution equipment and the image bounding box, and retrieve the corresponding three-dimensional control point model and type embedding vector from the database; The coordinate refinement module is used to concatenate the encoded frame image coordinates and the type embedding vector and input them into the conditional neural feature field to obtain scalar saliency values, and to perform dense sampling on the scalar saliency values within the image bounding box to obtain a saliency density map. The conditional neural feature field is constructed based on a fully connected network. Using the local maxima of the saliency density map as candidate control points, and performing gradient ascent iteration on the two-dimensional coordinates of each candidate control point, sub-pixel precision two-dimensional coordinates of each candidate control point are obtained. The joint prior module is used to obtain the truncated signed distance field of the region to be localized, and extract the depth value and gradient direction of the zero-level voxel set within the voxel region corresponding to the image bounding box to construct a joint prior for depth normals. The pose estimation module is used to construct a weighted constrained PnP objective function based on the three-dimensional control point model, the sub-pixel precision two-dimensional coordinates, and the joint prior of the depth normal, and solve it to obtain the estimated pose of the power distribution equipment.
[0013] Thirdly, embodiments of the present invention also provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method.
[0014] Fourthly, embodiments of the present invention also provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the above-described method.
[0015] This invention provides a subpixel precision positioning method, system, device, and storage medium for power distribution equipment. It replaces fixed feature operators with conditional neural feature fields, using a data-driven approach to uniformly solve subpixel positioning problems for various control points, thus eliminating feature type misjudgments. By combining the translational depth component and rotational normal direction of the PnP solution based on depth normals with prior constraints, it eliminates depth-rotation coupling ambiguity and improves the accuracy of pose estimation. Kalman filtering predicts feedforward motion compensation, enhancing the accuracy of two-dimensional coordinates. Incremental adaptive appearance updates ensure the system maintains high precision over long periods. This invention enables high-precision visual positioning for the precise control of power distribution equipment. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating the sub-pixel precision positioning method for power distribution equipment in an embodiment of the present invention; Figure 2 This is a schematic diagram of the sub-pixel precision positioning system for power distribution equipment in an embodiment of the present invention; Figure 3 This is an internal structural diagram of the computer device in an embodiment of the present invention.
[0017] Figure label: 10. Control point extraction module; 20. Coordinate refinement module; 30. Joint prior module; 40. Pose estimation module. Detailed Implementation
[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] Please see Figure 1 The first embodiment of the present invention proposes a sub-pixel precision positioning method for power distribution equipment, comprising steps S10 to S50: Step S10: Identify the frame images of the area to be located based on the quadruped robot to obtain the type of power distribution equipment and the image bounding box, and retrieve the corresponding three-dimensional control point model and type embedding vector from the database; Step S20: The encoded frame image coordinates and the type embedding vector are concatenated and input into the conditional neural feature field to obtain scalar saliency values. The scalar saliency values within the image bounding box are densely sampled to obtain a saliency density map. The conditional neural feature field is constructed based on a fully connected network. Step S30: Using the local maxima of the saliency density map as candidate control points, and performing gradient ascent iteration on the two-dimensional coordinates of each candidate control point to obtain the sub-pixel precision two-dimensional coordinates of each candidate control point. Step S40: Obtain the truncated signed distance field of the region to be localized, and extract the depth value and gradient direction of the zero-level voxel set within the voxel region corresponding to the image bounding box to construct a joint depth normal prior. Step S50: Based on the three-dimensional control point model, the sub-pixel precision two-dimensional coordinates, and the joint prior of the depth normal, a weighted constraint PnP objective function is constructed, and the estimated pose of the power distribution equipment is obtained by solving the function.
[0020] This embodiment provides a high-precision six-DOF visual positioning method for quadruped robots carrying robotic arms to perform precision control tasks on power distribution equipment. Based on the PnP subpixel positioning framework, it achieves subpixel control point positioning by fusing conditional neural feature fields, PnP solution of TSDF surface normal-depth joint constraints, Kalman predictive feedforward motion compensation, and incremental appearance adaptive updating. This enables sub-millimeter-level three-dimensional positioning of key targets such as distribution cabinet operating handles, circuit breaker buttons, and instrument panels in a power distribution substation environment without manual markings. It accurately guides the end effector of the robotic arm to complete precision control tasks such as pressing and turning, thereby meeting the actual needs of high-precision positioning in the automated operation and maintenance of power distribution networks.
[0021] This embodiment relies on a quadruped inspection robot platform. The preferred hardware configuration is as follows: a 1920×1080 global shutter industrial camera (30FPS frame rate, 2.4μm pixel size, 8mm focal length, supports hard triggering); a 16-line LiDAR (for SLAM mapping and TSDF maintenance); a MEMS-IMU (accelerometer range ±16g, sampling frequency 200Hz); triaxial force sensors at each leg end (range 0~500N, sampling frequency 1000Hz); and an onboard GPU (NVIDIA Jetson Orin NX, computing power ≥21TOPS). All sensors are synchronized via hardware trigger signals, with a synchronization error ≤0.5 ms. During the robot navigation phase, the SLAM module continuously updates the global 3D point cloud map and TSDF volume (voxel resolution 3cm). The robot's global pose is provided to the positioning module in real time, with a positioning accuracy error ≤2cm.
[0022] Based on the aforementioned hardware platform, this embodiment first determines the type of power distribution equipment by detecting power distribution equipment in the frame image of the current area. Specific steps include: A target detection network based on the YOLO model is used to identify targets in the frame images of the area to be identified, and the category labels, confidence scores and image bounding boxes of the power distribution equipment are obtained. Determine whether the confidence level is greater than the confidence threshold. If so, retrieve the three-dimensional control point model and type embedding vector from the preset database based on the category label.
[0023] In this embodiment, a lightweight object detection network on the onboard GPU is used to perform object recognition on the current frame image. Preferably, the lightweight object detection network is built using a YOLOv8 backbone network, and the architecture can be a standard YOLOv8 backbone network architecture. The lightweight object detection network outputs the category label, confidence score, and image bounding box for each power distribution device. To improve accuracy, this embodiment only performs subsequent positioning operations on power distribution devices with a confidence score greater than a confidence threshold, preferably set to 0.6.
[0024] This embodiment pre-constructs a geometric database of power distribution equipment within the current area. This database includes at least the following data types: a 3D control point model, comprising the coordinates of at least eight 3D control points in the equipment coordinate system; a type embedding vector, which is a semantic code trained offline using InfoNCE contrastive loss for each type of power distribution equipment, ensuring high cosine similarity for similar equipment and significant similarity for different types. In this embodiment, the type embedding vector is a 128-dimensional real vector. Furthermore, the geometric database also stores appearance drift thresholds (in pixels), which are empirical upper limits of the PnP reprojection residuals for various types of equipment under normal appearance, stored in tiers according to equipment type and years of service (e.g., a threshold of 1.2 pixels for a new XL-21 equipment and 1.8 pixels for equipment with 5 years of service); and equipment front normal vectors (e.g., standard values [0,0,-1]), control point types (e.g., center of distribution cabinet handle, corner of nameplate, center of opening, buttons, etc.), referenced GB / T standard numbers (e.g., GB / T 7251.1-2013), and nominal dimensions.
[0025] For power distribution equipment with a confidence level greater than the confidence threshold, the geometric database mentioned above is queried using the category label as the key to retrieve the 3D control point model and type embedding vector corresponding to the current equipment.
[0026] To address the insufficient robustness of fixed-type feature operators, this embodiment employs a conditional neural feature field to replace the fixed feature operator. This data-driven approach uniformly solves the sub-pixel localization problem for various control points, eliminates feature type misjudgments, and overcomes the matching instability of SIFT under varying lighting and worn appearance conditions. Specific steps include: Fourier position encoding is performed on the frame image coordinates, and the encoded frame image coordinates and the type embedding vector are concatenated and input into the conditional neural feature field to obtain scalar significance values; The coordinates within the image bounding box are densely sampled with a preset step size, and the scalar saliency values are queried in batches to obtain a saliency density map; The conditional neural feature field is constructed based on a multi-layer fully connected network, using a sine function as the activation function, and employing mean squared error loss based on focus weights and information noise contrast loss as loss functions.
[0027] In this embodiment, the conditional neural feature field preferably employs a 4-layer fully connected (MLP) network with a hidden layer width of 64 and uses sine activation, which naturally supports high-frequency function fitting. The input to the conditional neural feature field consists of two concatenated parts: one is the Fourier positional encoding of the image coordinates, and the other is the type embedding vector. Specifically, the image coordinates of the frame image are first Fourier positioned. Taking 128 dimensions as an example, the x-axis and y-axis coordinates are Fourier positioned separately and then concatenated to obtain 64 dimensions. This is then increased to 128 dimensions through a linear layer. This encoding can overcome the spectral bias of MLPs, enabling the network to learn sub-pixel level fine spatial patterns. The type embedding vector is also a 128-dimensional vector. The two are concatenated to obtain 256-dimensional input data. The concatenation of the type embedding vector and the Fourier positional encoding of the image coordinates is then input into the conditional neural feature field, allowing the same network to conditionally distinguish and process control points of different types of devices without requiring separate training for each device.
[0028] The output of the conditional neural feature field is a scalar significance value S∈[0,1]. This scalar significance value is the probability density of a certain type of control point at the input coordinates (x,y) (for example, S≈0.95 at the control point and S≈0.02 in the background region). S is continuously differentiable, allowing sub-pixel refinement by gradient ascent. A higher S value indicates that the point is closer to the control point.
[0029] Then, dense sampling is performed within the bounding box of the power distribution equipment image with a step size of 0.1 pixels. The global saliency density map S(x,y) is obtained through batch query. Non-maximum suppression (NMS) is performed on the saliency density map S(x,y) to extract local maxima as candidate control points. Since the input data of the above-mentioned conditional neural feature field includes type embedding vectors, the local maxima of its output scalar saliency value are consistent with the control points in the 3D control point model of the equipment type. Therefore, for different types of power distribution equipment, control points consistent with the control points in the 3D control point model can be screened by using the local maxima and the control point type corresponding to the scalar saliency value. For the sake of distinction, these are referred to as candidate control points.
[0030] Gradient ascent is performed with a preset step size and a maximum number of iterations. Preferably, the step size is set to 0.01 pixels and the maximum number of iterations is set to 20. The coordinates of the iterations can be represented as: In the formula, Let be the two-dimensional coordinates of the candidate control point at step t+1. Let be the two-dimensional coordinates of the candidate control point at step t. This is the gradient ascent step size. For the significance density plot in The two-dimensional gradient vector at image coordinate u.
[0031] Through the gradient ascent iteration described above, the coordinates of the candidate control points are refined to the maximum value of the saliency density map, thereby obtaining the sub-pixel precision two-dimensional coordinates of each candidate control point.
[0032] In a preferred embodiment, the conditional neural feature field is trained using a joint loss function, which includes a focal weight-based mean squared error loss (Focal MSE loss) and an information-noise contrast loss (InfoNCE loss). The focal weight-based mean squared error loss introduces focal weights on top of the standard MSE error. Since control points occupy only a very small number of positions (positive samples are sparse), the focal weights can prevent a large number of background negative sample gradients from dominating training. This loss function is mainly used to train scalar significance values, and its expression is as follows: In the formula, L fmse This represents the Focal MSE loss, where N is the total number of sampling points in the batch, and u d Let S(u) be the pixel coordinates of the d-th sampling point. d ) for network in u d The scalar significance prediction value at the output, G(u) d ) is a Gaussian soft label field in the same pixel u d The value at that location, The focus weight index is 2, and the preferred value is 2; the soft label field G(u d The definition of ) is: In the formula, For distance u d The true coordinates of the nearest control point, ||·|2 is the vector 2 norm, σ is the Gaussian kernel standard deviation, and the preferred value is 0.5 pixels.
[0033] The InfoNCE loss is used for joint training of type embeddings. Its purpose is to make the embeddings of similar devices have high cosine similarity and different types have high cosine similarity, thereby enhancing the type conditionation capability of the conditional neural feature field. Its expression is the regular InfoNCE loss function, which will not be elaborated here.
[0034] Total loss function It can be represented as: ,in, For the InfoNCE loss function, please refer to the general expression of the InfoNCE loss function; it will not be repeated here.
[0035] By using joint backpropagation, both the saliency prediction head and the type embedding encoder can be optimized simultaneously.
[0036] This embodiment learns the spatial location prior of control points in an end-to-end manner by conditionalizing the neural feature field, without relying on fixed texture description operators, thus eliminating the dependence on fixed appearance from a mechanistic perspective.
[0037] To address the depth-rotation coupling ambiguity issue in monocular point-to-point (PnP) scenarios with frontal view in corridors, this embodiment extracts depth and normal priors using a truncated signed distance function (TSDF) maintained by SLAM, constructing a two-dimensional joint prior constraint. The specific steps include: The truncated signed distance field of the region to be localized is obtained by SLAM technology, and the zero-level voxel set is extracted within the voxel region corresponding to the image bounding box. Extract the depth value set of each voxel in the zero-level voxel set in the camera coordinate system, and calculate the mean and standard deviation of the depth value set to obtain the depth prior mean and depth prior standard deviation. Extract the set of surface normal vectors from the zero-level collective element set, calculate the mean direction of the normal using the averaging method based on geodesic distance, and calculate the angular standard deviation of the angle between each normal and the mean direction of the normal to obtain the a prior mean direction of the normal and the a prior angular standard deviation of the normal.
[0038] This embodiment uses laser SLAM technology to obtain the truncated signed distance field (TSDF) of the current area. Laser SLAM is a technology that uses lidar data for real-time localization and map building. For the map building of a 3D scene, this embodiment uses a truncated signed distance field (TSDF) map. TSDF achieves incremental reconstruction of dynamic scenes by discretizing the 3D space into a voxel grid and recording the signed distance value of each voxel to the nearest surface. For the TSDF volume of lidar mapping, taking a voxel resolution of 3cm as an example, a zero-level voxel set (i.e., the device surface points measured by lidar) with an absolute TSDF value ≤ 6cm (2 voxels) within a radius of 0.5m centered on the target device center is extracted. Depth prior and normal prior are extracted from the zero-level voxel set.
[0039] For depth prior, including depth prior mean and depth prior standard deviation, the specific calculation steps include: extracting the set of z-coordinates (depth values) of each voxel in the zero-level voxel set in the camera coordinate system; for this set of depth values, calculating the depth value mean and standard deviation. The mean is used to characterize the average distance from the device surface to the camera, and the standard deviation is used to characterize the uncertainty of the depth distribution. The smaller the standard deviation, the tighter the depth constraint. The calculated mean and standard deviation are the depth prior mean and depth prior standard deviation.
[0040] In a preferred embodiment, to suppress outliers in lidar noise, both the depth prior mean and depth prior standard deviation are calculated using a weighted mean and weighted standard deviation method. The weights are calculated using the Huber weighting function, and the specific calculation steps include: 1) Initialization: Take the median of the sample as the initial mean. And the scale s is calculated based on the median absolute deviation (MAD): In the formula, Let be the a-th sample value (where, when calculating the depth prior, the depth value of the zero-level collective element in the camera coordinate system is taken, and when calculating the normal prior, the angle between each normal and the mean direction is taken), median(·) is the median, and 1.4826 is the consistency factor that makes MAD unbiased to the normal distribution.
[0041] 2) Calculate the standardized residual for the t-th iteration: In the formula, Let be the standardized residual of the a-th sample. Let be the sample mean of the t-th iteration.
[0042] 3) Calculate Huber weights: In the formula, Let a be the weight of the a-th sample. The Huber threshold is preferably set to 1.345.
[0043] 4) Update the mean: 5) Repeat steps 2) through 4) until the absolute value of the difference between two consecutive means is less than 10. -4 Or reach the maximum number of iterations (e.g., 10 times).
[0044] 6) Calculate the standard deviation: In the formula, denoted as the weighted standard deviation, and μ as the convergent weighted mean.
[0045] The mean and standard deviation obtained by the above weighting method are robust mean and standard deviation, which can effectively suppress the interference of outliers.
[0046] For the normal prior, including the normal prior mean direction and the normal prior angle standard deviation, the specific calculation steps are as follows: First, calculate the TSDF gradient direction for each zero-level ensemble in the zero-level ensemble set to obtain the surface normal vector set. Then, calculate the normal mean direction using the geodesic distance averaging method. The normal mean direction reflects the surface orientation of the device. Finally, calculate the angle between each normal vector and the normal mean, forming an angle set. Calculate the angle standard deviation of the angle set. The calculated normal mean direction and angle standard deviation are the normal prior mean direction and normal prior angle standard deviation. It should be noted that the geodesic distance averaging method is also called Karcher's mean, which is an average value calculated on the manifold. The specific calculation steps are the same as the conventional calculation steps for Karcher's mean, and will not be elaborated here.
[0047] The standard PnP objective function only minimizes the reprojection error, which introduces ambiguity due to depth rotation coupling (the objective function is flat and saddle-shaped when approaching from the front). To address this issue, this embodiment introduces a joint regularization term based on depth prior and normal prior to construct a weighted constrained PnP objective function. The specific steps include: Based on the three-dimensional control point model and the sub-pixel precision two-dimensional coordinates, a reprojection error term is constructed, wherein the coordinate point weights in the reprojection error term adopt the scalar significance values of the coordinate point positions. Based on the depth value of the translation vector and the prior mean of the depth, a depth regularization term is constructed, and based on the direction of the frontal normal vector of the device model coordinate system and the prior mean of the normal, a normal regularization term is constructed. The objective function of the weighted constraint PnP is to minimize the weighted sum of the reprojection error term, the depth regularization term, and the normal regularization term. The weight of the depth regularization term is determined based on the prior standard deviation of the depth, and the weight of the normal regularization term is determined based on the prior standard deviation of the normal angle.
[0048] In this embodiment, the weighted constraint PnP objective function is expressed as: In the formula, L(R,t) is the objective function value, R is a 3×3 rotation matrix, and t is a 3×1 translation vector. Together, they constitute the six-degree-of-freedom pose to be estimated [R|t]. i u is the significance weight of the i-th control point; iU represents the sub-pixel precision 2D coordinates of the i-th control point; π(·) is the perspective projection function used to project the 3D point onto the 2D pixel plane; K is the 3×3 camera intrinsic parameter matrix; U i Let t be the three-dimensional coordinates of the i-th control point in the device coordinate system; z Let z be the z-component of the translation vector t; The depth prior mean; These are the weighting coefficients for the depth regularization term; The weight coefficients for the normal regularization term; This is the normal vector to the front of the device model coordinate system. is the direction of the prior mean of the normal; the superscript T is the vector transpose; ‖·‖2 is the vector 2-norm.
[0049] The objective function described above consists of three parts: a reprojection error term, a depth regularization term, and a normal regularization term. The reprojection error term is the summation part of the objective function, representing the reprojection error between the two-dimensional and three-dimensional coordinates. U i Let w be the three-dimensional coordinates of the i-th control point in the device coordinate system. These coordinates are the three-dimensional coordinates of the control points contained in the three-dimensional coordinate model retrieved from the geometric database. i The saliency weight of the i-th control point is the output value of the saliency density map at the sub-pixel precision coordinates of the control point. For occluded / highly reflective regions S≈0, the weight value is approximately zero, contributing very little to the pose calculation. By designing the weight based on the scalar saliency value, automatic weight reduction can be achieved without introducing erroneous noise.
[0050] The middle part of the objective function is the depth regularization term, which is used to constrain the z-component of the translation vector t to not deviate from the prior depth mean. Here is the weighting coefficient for the depth regularization term. This coefficient is determined by the prior standard deviation of depth, and its calculation formula is: In the formula, This is the square of the standard deviation of the depth prior; the smaller the value, the more reliable the depth prior. The value is a proportionality coefficient, and the preferred value is 2.0.
[0051] The last part is the normal regularization term, used to constrain the rotation matrix R, so that the normal to the front of the device model coordinate system, after being rotated by R, is aligned with the direction of the prior mean of the normal. The weighting coefficient for the normal regularization term is determined by the standard deviation of the prior angle of the normal, and its calculation formula is as follows: In the formula, β is the square of the standard deviation of the prior angle of the normal. The smaller the value, the more reliable the prior angle of the normal. β is the proportionality coefficient, and the preferred value is 1.5.
[0052] The above objective function can eliminate the ambiguity of depth-rotation coupling. The principle is that when the view is close, the pure reprojection error is difficult to distinguish between a large depth and a small rotation angle and a small depth and a large rotation angle (the contour lines are all saddle-shaped). Here, "close view" refers to a scene where the angle between the camera's optical axis and the normal of the device's front elevation is very small, and the camera is approximately facing the front of the device. In this embodiment, an angle of 5° is considered a severe view scene. By using depth constraints, the z-class of the translation vector t is locked near the prior mean of depth, and the normal constraints limit R to a reasonable rotation range. After the two are combined, the contour lines are unimodal. The LM algorithm (Levenberg-Marquardt method) can stably converge to the unique correct solution. Therefore, for the above objective function, the LM algorithm is used to iteratively solve L(R,t) to output the accurate six-degree-of-freedom pose [R|t]. The specific solution steps are the conventional solution steps of the LM algorithm, which will not be repeated here.
[0053] In a preferred embodiment, to improve the stability of frame image capture, this embodiment provides an active frame image acquisition triggering mechanism, the specific steps of which include: Based on the triaxial force data and horizontal acceleration data of the quadruped robot's feet, a dynamic model of the quadruped robot based on body vibration was determined through data fitting. Based on the dynamic model, Kalman filtering is used to predict the planar displacement of the frame image to obtain the planar micro-displacement vector; The vibration amplitude is calculated based on the planar micro-displacement vector, and the vibration amplitude is compared with an amplitude threshold. In response to the vibration amplitude being less than or equal to the amplitude threshold, the quadruped robot is triggered to acquire frame images of the area to be located.
[0054] In this embodiment, before each positioning task begins (e.g., approximately 0.3 seconds before), the system is triggered to identify and collect the time series data from the foot triaxial force sensor (approximately 300 sampling points at 1000Hz) and the IMU horizontal acceleration, and then fit a second-order underdamped model using least squares. In the formula, The micro-vibration displacement of the inspection platform in the image plane direction (unit: pixel, which can be used to represent Δu in the x direction and Δv in the y direction, respectively). for The first derivative with respect to time; for The second derivative with respect to time, For the system damping ratio, The system's natural angular frequency, This is the external disturbance input, i.e., the equivalent input of the foot impact force.
[0055] By fitting the above second-order underdamped model, the natural frequency under the current ground conditions is identified. (Typical value approximately 2.0 Hz) and damping ratio ζ (typical value approximately 0.4) are used to determine the dynamic model of the quadruped robot under the current ground conditions, namely the second-order underdamped model. Based on this model, state sampling is performed at a sampling interval Δt = 1 / 200 s, and the state transition matrix F and observation matrix H are discretized to keep the model matching the current ground conditions.
[0056] Then, Kalman state initialization is performed, and the state vector is a 5-dimensional state vector. These represent the micro-displacement components in the x / y directions of the image plane, the first derivatives of the micro-displacement components in the x / y directions with respect to time (i.e., the micro-velocities in the x and y directions), the vibration reference acceleration state components, and the filter initial error covariance, respectively. `diag(·)` represents a diagonal matrix with the elements within the parentheses as diagonal elements and the rest as zero. The process noise covariance matrix Q and the observation noise covariance matrix R are also mentioned. obs Before model deployment, historical data is used to pre-set the parameters, following the conventional method for setting the process noise covariance matrix and observation noise covariance matrix, such as R... obs Q can be determined by the variance of the measured samples, such as by the variance of samples measured by the robot in a static environment. Q can be determined by the parameter tuning of the adaptive filter, such as by the time series of state variables (micro-displacement of the image plane and its rate of change) under typical gait using EM or autocovariance least squares method. The specific setting method can refer to the conventional setting method of process noise covariance matrix and observation noise covariance matrix, and is not limited here.
[0057] Then, a Kalman prediction update loop is performed, where: The prediction step is as follows: , ; The observations are: =[a IMUx ·scale, a IMUy ·scale,mean(Fz)·scale]ᵀ; The update steps are as follows: ; ; .
[0058] In the formula, and Let be the prior state estimate and the posterior state estimate at time k, respectively. Let F be the posterior state estimate at time k-1, and F be the state transition matrix. Let be the prior covariance matrix at time k. Let be the posterior covariance matrix at time k-1. Let be the posterior covariance matrix at time k. Let H be the Kalman gain at time k, and H be the observation matrix. Let Q be the observation noise covariance matrix, Q be the process noise covariance matrix, and I be the identity matrix. Let a be the observation vector. IMUx Let a be the IMU acceleration in the x-axis direction. IMUy y is the IMU acceleration in the y-axis direction, Fz is the foot force, mean(·) is the mean function, and scale is the preset calibration scale coefficient.
[0059] The estimated state vector at the current moment is obtained through Kalman filtering prediction, and the estimated micro-displacement component Δu in the x-direction is extracted from the vector. pred The estimated micro-displacement component Δv in the y-direction pred Based on the estimated micro-displacement component (Δu) pred ,Δv pred The system calculates the predicted vibration amplitude, which is the arithmetic square root of the sum of squares of the estimated micro-displacement components. Then, it determines whether the predicted vibration amplitude is greater than the amplitude threshold (e.g., 0.3 pixels). If not, it sends a hard trigger signal to the image acquisition camera for immediate exposure to ensure that the image is acquired at the vibration trough (the moment of minimum amplitude), thereby improving the stability of image acquisition and overcoming the inefficiency of passive stability weighting.
[0060] In a preferred embodiment, based on active acquisition triggering, this embodiment also provides a method for feedforward motion compensation of the detected coordinates based on Kalman filtering to further overcome the inefficiency of passive stability weighting. Specifically, after the current frame image is acquired, sub-pixel precision two-dimensional coordinates are obtained through conditional neural feature field detection and coordinate refinement. Then, the estimated micro-displacement components are used to perform feedforward correction on the sub-pixel precision two-dimensional coordinates. The corresponding estimated micro-displacement components are subtracted from the sub-pixel precision two-dimensional coordinates to obtain the corrected two-dimensional coordinates. Finally, the corrected two-dimensional coordinates are used to construct the PnP objective function. This embodiment, through feedforward correction based on Kalman filtering, can eliminate systematic image plane displacement errors introduced by vibration.
[0061] In the Kalman prediction feedforward motion compensation step, the final output, in addition to the corrected two-dimensional coordinates, also outputs the vibration confidence (within the range of [0,1]). The vibration confidence is linearly related to the predicted vibration amplitude. Taking the amplitude threshold range of [0.1,0.3] pixels as an example, when the predicted vibration amplitude is 0.1, the vibration confidence is 1, and when the predicted vibration amplitude is 0.3, the vibration confidence is 0.
[0062] Based on the dynamic model of body vibration, this embodiment establishes a Kalman filter to predict the micro-vibration state of the inspection platform in real time, which can realize feedforward motion compensation of detection coordinates and active acquisition triggering of frame images, thereby overcoming the inefficiency of passive stability weighting.
[0063] In a preferred embodiment, this embodiment employs an incremental appearance-adaptive higher-level mechanism to update the model, the specific steps of which include: Based on the estimated pose, calculate the reprojection residual of each 3D control point, and count the proportion of 3D control points whose reprojection residual exceeds the residual threshold. Based on the comparison between the ratio and the appearance drift threshold, it is determined whether appearance drift exists. The appearance drift threshold is retrieved from the database. If so, the parameters of the last layer of the conditional neural feature field are updated incrementally with microsteps driven by the reprojection residual gradient, and the 3D control points are rolled back and verified with geometric distance error as a hard constraint.
[0064] This embodiment establishes an online incremental learning closed loop for conditional neural feature fields. Before explaining the update steps, the key concepts are defined as follows: PnP convergence: The LM algorithm is considered convergent when iteratively solving the constrained objective function L(R,t) and satisfying any of the following conditions: ① The gradient norm of the objective function ② Pose change between two adjacent iterations ③ The number of iterations exceeds 100. After convergence, the optimal pose estimate [R|t] is output.
[0065] in, Let L be the gradient vector of the objective function L with respect to the pose parameters to be determined; ||·|2 is the vector 2 norm; ΔR and Δt are the changes in the rotation matrix and translation vector between two adjacent LM iterations, respectively; and These are the convergence criterion thresholds for the gradient norm and pose change, respectively. These thresholds are preferred values, but the specific thresholds can be flexibly set according to accuracy requirements.
[0066] Reprojection residual r j First, set the three-dimensional control point U jThe estimated pose [R|t] and camera intrinsic parameters K are projected onto the image: (Pinhole perspective projection); then calculate the reprojection residual: in, Let R be the predicted 2D pixel coordinates of the j-th control point after perspective projection of its estimated pose [R|t] and camera intrinsic parameters K, where R and t are the estimated rotation matrix and translation vector, respectively. Let j be the three-dimensional coordinates of the j-th control point in the equipment coordinate system. Let be the two-dimensional coordinates of the j-th control point after Kalman feedforward compensation, and ‖·‖2 be the vector 2-norm. The reprojection residual is the pixel distance between the predicted coordinates and the detected coordinates. A larger residual indicates a larger detection deviation for that control point, which may be caused by appearance drift. If no feedforward correction is performed, the reprojection residual is calculated using sub-pixel precision two-dimensional coordinates.
[0067] Based on the above key concept definitions, the steps for incremental adaptive appearance updates are as follows: Appearance drift trigger judgment: After PnP converges and the overall confidence is greater than the threshold (e.g., 0.7), the proportion of 3D control points whose reprojection residuals exceed the residual threshold is counted to the total number of control points; if the proportion is greater than the appearance drift threshold (e.g., 0.3), appearance drift is judged and the update process is entered; otherwise, it is skipped.
[0068] Update process: First, gradient calculation is performed, traversing all control points in the current frame that participate in the PnP solution, and then constructing an update loss function based on the reprojection residuals. ;r j For the reprojection residual of the j-th control point, backpropagation is performed on the parameters of the last two layers of the conditional neural feature field (with the gradients of the remaining layers set to zero) according to the update loss, and the gradient is calculated. , This is the set of learnable parameters for the last two layers of the conditional neural feature field.
[0069] Then, the Adam optimizer is used for micro-step updates: θ tail ←Adam(θ tail lr=1×10 -4 Only one update is performed to prevent overfitting, where lr represents the learning rate.
[0070] Finally, a geometric constraint rollback check is performed. After the model update, the geometric distance error increment is recalculated for all control point pairs. The geometric distance error increment is the change in GB geometric distance error before and after the last layer incremental update. If the geometric distance error increment of any control point pair exceeds 0.5 pixels, a rollback is performed. The updated value is discarded to reflect the previous value to prevent the network from deviating from the correct geometry. This includes the geometric distance error increment. The calculation is as follows: For any control point pair (i,j), the reprojected pixel distance is obtained by projecting the poses before and after the update respectively: Then the geometric distance error increment of the control point pair (i,j) for: In the formula, and Let (i,j) be the two-dimensional projected pixel coordinates of the control point pair (i,j) before the final incremental update. and Let (i,j) be the two-dimensional projected pixel coordinates of the control point pair (i,j) after the final incremental update. and , respectively, are the reprojected pixel distances of the control point pair (i,j) before and after the last layer incremental update, ‖·‖2 is the vector 2 norm, and |·| is the absolute value operation.
[0071] Finally, the differential weights are persisted, and after the update passes the verification, they are... The differential weight file of the device instance is incrementally written and automatically loaded the next time a similar device is located, enabling continuous adaptation across tasks without the need for manual recalibration.
[0072] This embodiment establishes an online incremental learning closed loop for conditional neural feature fields, enabling the system to automatically adapt to device appearance drift during long-term operation, maintain high-precision positioning capabilities, and eliminate the need for manual recalibration.
[0073] In a preferred embodiment, precise control of power distribution equipment is achieved through standard geometric constraint rationality verification and a three-level decision-making mechanism. Specific steps include: All control points are paired into control point pairs. The geometric consistency error is calculated for all control point pairs. The geometric consistency error is the absolute value of the difference between the measured reprojection image distance of the control point pair and the standard perspective prediction distance. The geometric consistency score is calculated based on the geometric consistency error, the reprojection accuracy score is calculated based on the reprojection residual, and the overall confidence score is calculated based on the geometric consistency score, the reprojection accuracy score, and the vibration confidence score. Based on the comparison between the comprehensive confidence score and the scoring threshold range, the reliability of the estimated pose is determined, and the corresponding decision is made based on the reliability.
[0074] In this embodiment, the geometric consistency error is first calculated. Specifically, after PnP convergence, any two control points are considered as a pair, and the measured reprojection image distance and the standard perspective prediction distance are calculated for all control point pairs. The measured reprojection image distance is the reprojection pixel distance of the two-dimensional observation coordinates of the control point pair based on the actual image detection and after sub-pixel refinement and feedforward compensation. The standard perspective prediction distance is the pixel distance of the predicted two-dimensional pixel coordinates obtained based on the current estimated pose and the camera intrinsic parameters of the forward perspective projection. Then, the absolute value of the difference between the measured reprojection image distance and the standard perspective prediction distance is taken as the geometric consistency error.
[0075] The geometric consistency score is calculated by subtracting this ratio from the total number of control point pairs whose geometric consistency error exceeds a threshold (e.g., 0.1). In the formula, For geometric consistency score, Let be the geometric consistency error of the control point pair consisting of control point j and control point m. To control the total number of point pairs, count(·) represents the counting function.
[0076] The reprojection residuals of all control points are calculated according to the calculation steps of the above embodiment, and the reprojection accuracy score is calculated using an exponential function: In the formula, For the reprojection accuracy score, exp(·) is the natural exponential function, mean(·) represents the mean function, and r j σ represents the reprojection residual of control point j. r This is the reprojection accuracy normalization constant, with a value of 1.0.
[0077] Finally, the geometric mean of the geometric consistency score, reprojection accuracy score, and vibration confidence score is calculated as the overall confidence score: In the formula, To calculate the overall confidence score, For vibration confidence level.
[0078] A three-level decision is made based on the comprehensive confidence score. Taking a score threshold range of [0.52, 0.72) as an example, when the comprehensive confidence score is within this range, the pose estimation is considered uncertain, and the localization process is retried (images are reacquired); if it exceeds the upper limit of the score range, the estimated pose is considered reliable, [R|t] is output, and the robotic arm control command is triggered. The coordinates are mapped to the robot base coordinate system through SLAM coordinate transformation, thereby guiding the robotic arm end effector to perform precision control tasks; if it is below the lower limit of the score range, the localization is considered to have failed, an alarm is triggered, and manual confirmation is required or the robot is readjusted before trying again.
[0079] The positioning effect of this embodiment is verified by scene positioning test. Taking the precise positioning of the handle of the XL-21 low-voltage distribution cabinet in a frontal approach scenario as an example, in this scenario, in the corridor of a 10 kV substation, the robot stops 1.3m away from the XL-21 distribution cabinet with an approach angle θ=5° (severe frontal approach, the depth rotational coupling of PnP is most significant). The positioning is performed by the positioning method provided in this embodiment, and the comprehensive confidence score is calculated to be 0.94. The control is triggered. The test results are: the three-dimensional positioning error of the handle center is 0.35mm (x / y / z: 0.21 / 0.19 / 0.22mm). The robotic arm control is successful. Among them, 0.35mm is the comprehensive three-dimensional positioning error (three-dimensional Euclidean distance) between the estimated three-dimensional coordinates of the handle center and the true value. The 0.21 / 0.19 / 0.22mm in parentheses are the component errors along the X, Y and Z axes, respectively.
[0080] The comparative verification showed that removing the normal constraint (retaining only the depth constraint) increased the rotation angle deviation to 1.8° and the end effector position deviation to 3.1 mm, resulting in control failure. After restoring the normal constraint, the deviation decreased to 0.35 mm. This experimental result verifies the key correction role of the normal constraint on the rotation component when approaching from the front, which also verifies the positioning accuracy of the positioning method provided in this embodiment.
[0081] This embodiment provides a subpixel precision positioning method for power distribution equipment. By replacing fixed feature operators with conditional neural feature fields, it solves the subpixel positioning problem of various control points in a unified data-driven manner, eliminating feature type misjudgments. By combining the translational depth component and rotational normal direction of the PnP solution based on the depth normal, it eliminates depth-rotation coupling ambiguity and improves the accuracy of pose estimation. Kalman filtering predicts feedforward motion compensation to improve the accuracy of two-dimensional coordinates. Incremental adaptive appearance updates enable the system to maintain high precision over a long period. The method provided in this embodiment can achieve high-precision visual positioning for the precise control of power distribution equipment.
[0082] Please see Figure 2 Based on the same inventive concept, the second embodiment of the present invention proposes a sub-pixel precision positioning system for power distribution equipment, comprising: The control point extraction module 10 is used to identify the frame image of the area to be located based on the quadruped robot, obtain the type of power distribution equipment and the image bounding box, and retrieve the corresponding three-dimensional control point model and type embedding vector from the database. The coordinate refinement module 20 is used to concatenate the encoded frame image coordinates and the type embedding vector and input them into the conditional neural feature field to obtain scalar saliency values, and to perform dense sampling on the scalar saliency values within the image bounding box to obtain a saliency density map. The conditional neural feature field is constructed based on a fully connected network. Using the local maxima of the saliency density map as candidate control points, and performing gradient ascent iteration on the two-dimensional coordinates of each candidate control point, sub-pixel precision two-dimensional coordinates of each candidate control point are obtained. The joint prior module 30 is used to obtain the truncated signed distance field of the region to be localized, and extract the depth value and gradient direction of the zero-level voxel set within the voxel region corresponding to the image bounding box to construct a joint prior for the depth normal. The pose estimation module 40 is used to construct a weighted constrained PnP objective function based on the three-dimensional control point model, the sub-pixel precision two-dimensional coordinates and the joint prior of the depth normal, and solve it to obtain the estimated pose of the power distribution equipment.
[0083] The technical features and effects of the sub-pixel precision positioning system for power distribution equipment proposed in this embodiment are the same as those of the method proposed in this embodiment, and will not be repeated here. Each module in the above-mentioned sub-pixel precision positioning system for power distribution equipment can be implemented entirely or partially through software, hardware, or a combination thereof. Each module can be embedded in or independent of the processor in a computer device in hardware form, or it can be stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0084] Furthermore, embodiments of the present invention also propose a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described method.
[0085] Please see Figure 3The diagram illustrates the internal structure of a computer device in one embodiment. This computer device can specifically be a terminal or a server. The computer device includes a processor, memory, network interface, display, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the computer device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements a sub-pixel precision positioning method for power distribution equipment. The display screen of the computer device can be a liquid crystal display (LCD) or an e-ink display. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0086] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computing devices may include more or fewer components than those shown in the figure, or combine certain components, or have the same component arrangement.
[0087] Furthermore, embodiments of the present invention also propose a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the above-described method.
[0088] In summary, the embodiments of this invention propose a sub-pixel precision positioning method, system, device, and storage medium for power distribution equipment. The method identifies the type and bounding box of the power distribution equipment by analyzing frame images of the area to be positioned acquired by a quadruped robot, and retrieves the corresponding 3D control point model and type embedding vector from a database. The encoded frame image coordinates and the type embedding vector are concatenated and input into a conditional neural feature field to obtain scalar saliency values. Dense sampling is then performed on the scalar saliency values within the bounding box to obtain a saliency density map. The conditional neural feature field is based on a full... The network is constructed; local maxima of the saliency density map are used as candidate control points, and gradient ascent iteration is performed on the two-dimensional coordinates of each candidate control point to obtain sub-pixel precision two-dimensional coordinates; the truncated signed distance field of the region to be located is obtained, and the depth value and gradient direction of the zero-level voxel set are extracted in the voxel region corresponding to the image bounding box to construct a joint depth-normal prior; based on the three-dimensional control point model, the sub-pixel precision two-dimensional coordinates and the joint depth-normal prior, a weighted constrained PnP objective function is constructed, and the estimated pose of the power distribution equipment is obtained by solving it. This invention replaces fixed feature operators with conditional neural feature fields, providing a unified data-driven solution for sub-pixel localization of various control points and eliminating feature type misjudgments. By combining the translational depth component and rotational normal direction of the PnP solution with the depth normal, it eliminates depth-rotation coupling ambiguity and improves the accuracy of pose estimation. Kalman filtering predicts feedforward motion compensation, enhancing the accuracy of two-dimensional coordinates. Incremental adaptive appearance updates ensure the system maintains high precision over the long term. This invention enables high-precision visual positioning for the precise control of power distribution equipment.
[0089] The various embodiments in this specification are described in a progressive manner. For directly identical or similar parts of the embodiments, refer to each other. Each embodiment focuses on its differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments. It should be noted that the technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0090] The embodiments described above are merely preferred embodiments of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various improvements and substitutions without departing from the technical principles of this invention, and these improvements and substitutions should also be considered within the scope of protection of this application. Therefore, the scope of protection of this patent application should be determined by the scope of the claims.
Claims
1. A sub-pixel precision positioning method for power distribution equipment, characterized in that, include: The frame images of the area to be localized, collected by a quadruped robot, are identified to obtain the type of power distribution equipment and the image bounding box. The corresponding three-dimensional control point model and type embedding vector are retrieved from the database. The encoded frame image coordinates and the type embedding vector are concatenated and input into the conditional neural feature field to obtain scalar saliency values. The scalar saliency values within the image bounding box are then densely sampled to obtain a saliency density map. The conditional neural feature field is constructed based on a fully connected network. Using the local maxima of the saliency density map as candidate control points, and performing gradient ascent iteration on the two-dimensional coordinates of each candidate control point, sub-pixel precision two-dimensional coordinates of each candidate control point are obtained. Obtain the truncated signed distance field of the region to be localized, and extract the depth value and gradient direction of the zero-level voxel set within the voxel region corresponding to the image bounding box to construct a joint depth normal prior; Based on the three-dimensional control point model, the sub-pixel precision two-dimensional coordinates, and the joint prior of the depth normal, a weighted constrained PnP objective function is constructed, and the estimated pose of the power distribution equipment is obtained by solving it.
2. The sub-pixel precision positioning method for power distribution equipment according to claim 1, characterized in that, The steps of concatenating the encoded frame image coordinates and the type embedding vector and inputting them into the conditional neural feature field to obtain scalar saliency values, and then densely sampling the scalar saliency values within the image bounding box to obtain a saliency density map include: Fourier position encoding is performed on the frame image coordinates, and the encoded frame image coordinates and the type embedding vector are concatenated and input into the conditional neural feature field to obtain scalar significance values; The coordinates within the image bounding box are densely sampled with a preset step size, and the scalar saliency values are queried in batches to obtain a saliency density map; The conditional neural feature field is constructed based on a multi-layer fully connected network, using a sine function as the activation function, and employing mean squared error loss based on focus weights and information noise contrast loss as loss functions.
3. The sub-pixel precision positioning method for power distribution equipment according to claim 1, characterized in that, The steps of obtaining the truncated signed distance field of the region to be localized, and extracting the depth value and gradient direction of the zero-level voxel set within the voxel region corresponding to the image bounding box, and constructing a joint depth-normal prior include: The truncated signed distance field of the region to be localized is obtained by SLAM technology, and the zero-level voxel set is extracted within the voxel region corresponding to the image bounding box. Extract the depth value set of each voxel in the zero-level voxel set in the camera coordinate system, and calculate the mean and standard deviation of the depth value set to obtain the depth prior mean and depth prior standard deviation. Extract the set of surface normal vectors from the zero-level collective element set, calculate the mean direction of the normal using the averaging method based on geodesic distance, and calculate the angular standard deviation of the angle between each normal and the mean direction of the normal to obtain the a prior mean direction of the normal and the a prior angular standard deviation of the normal.
4. The sub-pixel precision positioning method for power distribution equipment according to claim 3, characterized in that, The step of constructing a weighted constrained PnP objective function based on the three-dimensional control point model, the sub-pixel precision two-dimensional coordinates, and the joint prior of the depth normal includes: Based on the three-dimensional control point model and the sub-pixel precision two-dimensional coordinates, a reprojection error term is constructed, wherein the coordinate point weights in the reprojection error term adopt the scalar significance values of the coordinate point positions. Based on the depth value of the translation vector and the prior mean of the depth, a depth regularization term is constructed, and based on the direction of the frontal normal vector of the device model coordinate system and the prior mean of the normal, a normal regularization term is constructed. The objective function of the weighted constraint PnP is to minimize the weighted sum of the reprojection error term, the depth regularization term, and the normal regularization term. The weight of the depth regularization term is determined based on the prior standard deviation of the depth, and the weight of the normal regularization term is determined based on the prior standard deviation of the normal angle.
5. The sub-pixel precision positioning method for power distribution equipment according to claim 1, characterized in that, Before the step of recognizing the frame images of the region to be localized acquired by the quadruped robot, the method further includes: Based on the triaxial force data and horizontal acceleration data of the quadruped robot's feet, a dynamic model of the quadruped robot based on body vibration was determined through data fitting. Based on the dynamic model, Kalman filtering is used to predict the planar displacement of the frame image to obtain the planar micro-displacement vector; The vibration amplitude is calculated based on the planar micro-displacement vector, and the vibration amplitude is compared with an amplitude threshold. In response to the vibration amplitude being less than or equal to the amplitude threshold, the quadruped robot is triggered to acquire frame images of the area to be located.
6. The sub-pixel precision positioning method for power distribution equipment according to claim 5, characterized in that, After the step of obtaining the sub-pixel precision two-dimensional coordinates of each candidate control point, the method further includes: The sub-pixel precision two-dimensional coordinates of each candidate control point are feedforward corrected based on the planar micro-displacement vector to obtain the corrected two-dimensional coordinates.
7. The sub-pixel precision positioning method for power distribution equipment according to claim 1, characterized in that, After the step of calculating the estimated pose of the power distribution equipment, the method further includes: Based on the estimated pose, calculate the reprojection residual of each 3D control point, and count the proportion of 3D control points whose reprojection residual exceeds the residual threshold. Based on the comparison between the ratio and the appearance drift threshold, it is determined whether appearance drift exists. The appearance drift threshold is retrieved from the database. If so, the parameters of the last layer of the conditional neural feature field are updated incrementally with microsteps driven by the reprojection residual gradient, and the 3D control points are rolled back and verified with geometric distance error as a hard constraint.
8. A sub-pixel precision positioning system for power distribution equipment, characterized in that, include: The control point extraction module is used to identify the frame images of the area to be localized based on the quadruped robot, obtain the type of power distribution equipment and the image bounding box, and retrieve the corresponding three-dimensional control point model and type embedding vector from the database; The coordinate refinement module is used to concatenate the encoded frame image coordinates and the type embedding vector and input them into the conditional neural feature field to obtain scalar saliency values, and to perform dense sampling on the scalar saliency values within the image bounding box to obtain a saliency density map. The conditional neural feature field is constructed based on a fully connected network. Using the local maxima of the saliency density map as candidate control points, and performing gradient ascent iteration on the two-dimensional coordinates of each candidate control point, sub-pixel precision two-dimensional coordinates of each candidate control point are obtained. The joint prior module is used to obtain the truncated signed distance field of the region to be localized, and extract the depth value and gradient direction of the zero-level voxel set within the voxel region corresponding to the image bounding box to construct a joint prior for depth normals. The pose estimation module is used to construct a weighted constrained PnP objective function based on the three-dimensional control point model, the sub-pixel precision two-dimensional coordinates, and the joint prior of the depth normal, and solve it to obtain the estimated pose of the power distribution equipment.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.