Simulation Method for Background Replacement in Wiring Images of Power Grid Robots Based on Dual-View Consistency Constraints

By combining dual-view consistency constraints and the U-Net segmentation network, the view consistency problem of background replacement in the multi-view camera system of the power grid robot is solved, generating a high-fidelity and diverse simulation dataset that meets the accuracy requirements of visual model training.

CN122200242BActive Publication Date: 2026-07-17NANJING UNIV OF AERONAUTICS & ASTRONAUTICS

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING UNIV OF AERONAUTICS & ASTRONAUTICS
Filing Date
2026-05-15
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Traditional background replacement techniques cannot meet the multi-view collaborative requirements of multi-camera systems for power grid robots, resulting in logical inconsistencies in cross-view background space and contradictory parallax relationships. Consequently, the generated simulation data cannot meet the accuracy requirements for visual model training.

Method used

A simulation method for changing the background of a power grid robot wiring image based on dual-view consistency constraints is adopted. An integrated framework is constructed by combining view consistency constraints, correlation-based new background generation, semantic segmentation, foreground and new background synthesis and optimization techniques. The foreground mask of the power line-robot system is extracted through the U-Net segmentation network, and the white edge jaggedness and color lighting of the foreground and background are optimized.

Benefits of technology

It achieves consistent background replacement from dual perspectives, constructs a high-fidelity and diverse simulation dataset for power grid robot operations, and improves the accuracy of visual model training and its ability to adapt to complex outdoor scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122200242B_ABST
    Figure CN122200242B_ABST
Patent Text Reader

Abstract

This invention proposes a simulation method for background replacement in wiring images of power grid robots based on dual-view consistency constraints. Relating to the fields of image processing and power grid robot technology, the method first acquires wiring operation scene images collected by the fisheye binocular vision system of the power grid robot. A cross-view correlation mapping model is constructed based on the consistency constraint relationship between the left and right dual views. Then, a new background with this view relationship is obtained by combining the background image to be replaced with the new background image. The U-Net segmentation network is used to extract the foreground mask of the power line-robot system. Finally, the foreground images from the left and right views are synthesized with the new background image, and the jagged edges and color / lighting inconsistencies of the synthesized image are optimized. This invention performs view-consistency background replacement on wiring operation images from the left and right views of the power grid robot, fully preserving the left-right view correlation in the generation of new backgrounds, and providing diverse dataset support for multi-view visual images of power grid robots.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of image processing and power grid robot technology, specifically a simulation method for changing the background of wiring images for power grid robots based on dual-view consistency constraints. Background Technology

[0002] With the deepening of smart grid construction, power grid operation robots have become core equipment for improving the safety and efficiency of power operation and maintenance. The multi-view camera vision system they carry is the core module for realizing environmental perception and operation control. The training of high-performance vision algorithms relies heavily on large-scale, multi-scene image datasets. However, training data with a single fixed background is prone to causing background bias in the model, resulting in insufficient generalization ability and difficulty in adapting to the complex and ever-changing outdoor scenes in real-world operations.

[0003] Traditional background replacement techniques are mostly designed for processing single images independently, and can only replace backgrounds from a single perspective. This makes them unsuitable for the multi-view collaborative requirements of multi-camera systems in power grid robots. Performing independent background replacement on multi-view images can lead to cross-view background spatial logic disorder, contradictory parallax relationships, and damage to the inherent geometric consistency of multiple perspectives. Consequently, the generated simulation data cannot meet the accuracy requirements for training visual models.

[0004] This invention addresses the problems of lack of cross-view geometric consistency and inability of synthetic data to meet the training requirements of visual models in traditional single-view independent background replacement of fisheye binocular camera systems for power grid robots. In order to achieve synchronous background replacement from two perspectives and construct a high-fidelity and diverse power grid operation simulation dataset, this invention proposes a simulation method for background replacement of wiring images for power grid robots based on dual-view consistency constraints. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention aims to provide a simulation method for background replacement in power grid robot wiring images based on dual-view consistency constraints. This method integrates view consistency constraints, associative new background generation, semantic segmentation, foreground and new background synthesis, and optimization techniques to construct an integrated framework. For wiring operation scene images acquired by the fisheye binocular vision system of the power grid robot, a cross-view correlation mapping model is first constructed through the consistency constraint relationship between the left and right dual views. Then, a new background with this view relationship is obtained by combining it with the background image to be replaced. The foreground mask of the power line-robot system is extracted using a U-Net segmentation network. Finally, the foreground images from the left and right views are synthesized with the new background image, and the jagged edges of the foreground and background and the color and lighting inconsistencies in the synthesized image are optimized. Compared with single-view independent background replacement methods, this invention can achieve dual-view consistent background replacement and construct a high-fidelity and diverse power grid robot operation simulation dataset.

[0006] This application achieves the above-mentioned effect through the following technical solution: a simulation method for changing the background of wiring images of a power grid robot based on dual-view consistency constraints, characterized in that the method includes the following steps:

[0007] S1. Obtain the left and right view wiring operation scene image sequences Image1_left and Image1_right collected by the fisheye binocular vision system of the power grid robot, and construct a power grid robot wiring image sample library based on the image sequences;

[0008] S2 performs camera calibration on the image sequences Image1_left and Image1_right respectively, obtains the intrinsic parameters and distortion coefficients Data1_left and Data1_right of the left and right cameras respectively, calculates the consistency constraint relationship between the left and right dual views, and constructs the cross-view correlation mapping model Model1.

[0009] S3 unifies the background image Image2 to be replaced into a fisheye type image, then performs image denoising to obtain Image2_ready, and then maps Image2_ready and the image sequences Image1_left and Image1_right from the fisheye pixel plane to the unit sphere respectively;

[0010] S4 constructs a mapping relationship between Image2_ready and the target fisheye new background image through the unit spherical coordinate system, generates new backgrounds from the left and right perspectives, and obtains the visualized new background images Image3_left and Image3_right from the left and right perspectives.

[0011] S5 uses the U-Net segmentation network to perform semantic segmentation on the image sequences Image1_left and Image1_right, and extracts the foreground mask images Image4_left and Image4_right of the power line-robot system.

[0012] S6 The image sequences Image1_left and Image1_right, the new background images Image3_left and Image3_right from the left and right perspectives, and the foreground mask images Image4_left and Image4_right are respectively synthesized by combining the foreground images and new background images from the left and right perspectives to obtain the synthesized images Image5_left and Image5_right from the left and right perspectives.

[0013] S7 optimizes the foreground and background edges with jagged white edges and color and lighting inconsistencies in the synthesized images Image5_left and Image5_right, resulting in the final optimized images Image6_left and Image6_right.

[0014] Furthermore, S2 specifically includes:

[0015] S21. Based on a single-image wiring operation scenario, a Unified Spherical Model (USM) parameter prediction method is used to construct a multi-task prediction model with EfficientNetv2 as the backbone network. Given a wiring operation scenario image... The output of the prediction model network is represented as:

[0016] ;

[0017] in, , , and For continuous regression values, they correspond to the estimated values ​​of the field of view of the fisheye camera of the power grid robot, the estimated values ​​of the USM nonlinear distortion parameters, and the estimated values ​​of the pitch and roll angles, respectively. The pitch angle category prediction results are used to assist in the stable training of attitude estimation;

[0018] S22 constructs the focal length of the fisheye camera for the power grid robot through analytical calculation within the framework of the USM parameter prediction method. , Images of wiring operation scenarios width The field of view and USM nonlinear distortion parameters are used to represent:

[0019] ;

[0020] Image of wiring operation scene The principal point is taken as the image center. In the formula, , These are images of the wiring operation scene. The width and height are thus determined, allowing for assembly. Intrinsic parameter matrix:

[0021] ;

[0022] The output of the prediction model network in S21 After inverse normalization and scale recovery, the USM nonlinear distortion parameters in the true physical dimensions are obtained. Perform the construction evolution steps S21 to S22 on the image sequences Image1_left and Image1_right respectively to obtain the intrinsic parameter matrices of the left and right cameras. , and their respective distortion coefficients , This yields Data1_left and Data1_right;

[0023] S23 takes the pitch angle estimate output by the prediction model network in S21. Compared with the estimated roll angle After inverse normalization and scale recovery, the pitch angle in true physical dimensions is obtained. With roll angle The yaw angle is manually set according to the specific application scenario. ,according to - - Constructing rotation matrices in Euler angle order: Complete the attitude description of the camera in the world coordinate system;

[0024] S24 models the binocular vision system as a common pose perturbation under a rigid body structure and models the background image to be replaced as a spherical environment texture located at infinity.

[0025] S25 sets the left and right cameras to have basic yaw correction values ​​respectively. and This describes the fixed deviation of each installation orientation relative to a reference direction. For the background image to be replaced, this is used to obtain the first... A geometrically consistent background simulation sample is used, with a yaw disturbance shared by the left and right cameras applied. The rotation matrices of the left and right cameras are defined as follows:

[0026] ;

[0027] in, This represents the rotation matrix about the vertical axis. Representing the left or right camera, a cross-view correlation mapping model Model1 is constructed.

[0028] Furthermore, S3 specifically includes:

[0029] S31 sets the background image to be replaced. Image2 supports two image types: ERP panoramic image and standard pinhole image. The ERP panoramic image itself is captured by a fisheye camera and is a fisheye type image.

[0030] For the standard pinhole image type background image to be replaced, the isometric projection model commonly used in industrial fisheye cameras is used to convert it into a fisheye type image;

[0031] S32 uses bilateral filtering on the fisheye-type background image to be replaced. Apply edge-preserving noise reduction processing: for pixels at any position The denoising result is:

[0032] ;

[0033] in, For pixels The local neighborhood, For the neighborhood Any pixel within, and The smoothing intensity is controlled separately in the spatial domain and the grayscale domain. The normalization coefficient is... The denoised image is denoised as Image2_ready; the denoised ERP panoramic image is denoised as... A standard pinhole image, after fisheye conversion and denoising, is denoised as... ;

[0034] S33 for images Set its resolution to For any three-dimensional direction vector in the unit spherical coordinate system ,satisfy ,Will Convert to spherical angular coordinates:

[0035] ;

[0036] ;

[0037] in, for azimuth angle, for The pitch angle; where, To maintain consistency with the downward direction convention of the vertical axis in the ERP panoramic view coordinate system; then... Mapping from spherical angular coordinates to ERP 2D pixel texture coordinates:

[0038] ;

[0039] This establishes a mapping relationship between unit spherical angular coordinates and ERP two-dimensional pixel texture coordinates, based on which... Image pixels are desampled to a unit sphere;

[0040] S34 for images For any three-dimensional direction vector in the same unit spherical coordinate system , exist The normalized coordinates on the image plane are: Thus, a unit spherical coordinate system is constructed. The mapping relationship of two-dimensional pixel texture coordinates in an image, Regions whose projection points extend beyond the image boundaries are considered invalid regions; for valid regions, [the following is omitted as the text is incomplete and cannot be translated]. Image pixels are desampled to a unit sphere;

[0041] For wiring operation scenarios, the S35 inputs image pixel coordinates under a unified spherical model. Solve for the three-dimensional direction vector in the unit spherical coordinate system. :

[0042] ;

[0043] ;

[0044] ;

[0045] in, , , , Principal point of the image, For the focal length of the fisheye camera of the power grid robot, The nonlinear distortion parameters of the USM are used; the mapping from image pixel coordinates to unit spherical coordinates is denoted as the inverse mapping. Perform inverse mapping on the image sequences Image1_left and Image1_right respectively. The images of the wiring operation scene from both the left and right perspectives are projected from the fisheye pixel plane onto the unit sphere.

[0046] Furthermore, S4 specifically includes:

[0047] S41 introduces a focal length scaling factor Intrinsic parameter matrix of fisheye camera for power grid robot Perform proportional scaling. Indicates the left or right camera, and obtains the adjusted intrinsic parameter matrix. :

[0048] ;

[0049] in, For construction The camera's focal length for Camera image principal point coordinates; Corresponding to the magnified focal length, Correspondingly, reduce the focal length. Then the original focal length remains unchanged;

[0050] S42 wiring operation scene image Resolution size is In order to obtain the same The target fisheye background images Image3_left and Image3_right are defined with the following pixel domains: ;

[0051] S43 Any pixel coordinate in the target fisheye background image Using the reverse mapping established in S35 Based on the parameters of the fisheye camera of the power grid robot Map it to a three-dimensional direction vector in the unit spherical coordinate system: ;in, Represents the unit sphere, and indicates the output vector. The unit line-of-sight direction vector in three-dimensional space;

[0052] S44 uses the rotation matrix constructed in S25. Apply a viewpoint rotation transformation, Transform the background texture orientation onto the unit sphere to obtain the rotated orientation vector: ;

[0053] S45, based on the source of the background image on the unit sphere, rotates the direction vector... Through different sampling mappings Convert to background image or The two-dimensional texture coordinates in the image; the background image on the unit sphere is derived from... hour, Calculated from S33; derived from hour, Calculated from S34;

[0054] S46 uses bilinear interpolation from the background image or The pixel values ​​are obtained to generate the new background image for the target fisheye.

[0055] S47 Different pixel coordinates of the target fisheye background image The pixel value at that location is given by the following formula:

[0056] ;

[0057] in, In the first Under a background simulation sample with consistent geometry from each viewpoint The camera's target fisheye background image; Image2_ready refers to the background image. or This results in two new background images, Image3_left and Image3_right, visualized from the left and right perspectives.

[0058] Furthermore, S5 specifically includes:

[0059] S51 Constructs the power grid robot wiring image dataset Image1_dataset based on the image sequences Image1_left and Image1_right;

[0060] S52 Use image annotation software to annotate the power line-robot system region of a single image in the image dataset Image1_dataset, set it as a label to represent the category of the foreground region, save the annotation file in the same path as the image, and generate a JSON format annotation file containing the foreground region category;

[0061] S53 The labeled files are batch converted into binary mask format files required for training, and a semantic segmentation labeled dataset Labeling1_dataset containing original image-mask pairs is constructed.

[0062] S54. The semantic segmentation annotation dataset Labeling1_dataset is divided into training and testing sets. A U-Net network based on an encoder-decoder structure is constructed, and a focus loss function is built. and Dice coefficient loss Weighted Mixed Loss Function The expression is:

[0063] ;

[0064] ;

[0065] in, To balance the weighting coefficients, For the first The predicted probability value of each pixel. No. The actual label value of each pixel. Total number of pixels For smoothing terms;

[0066] S55 Iteratively trains the U-Net network using the training set and the hybrid loss function. When the Intersection over Union (IoU) index of the test set reaches a preset threshold, the optimal model weights are saved. The optimal model weights are loaded to perform inference on the image sequences Image1_left and Image1_right, and output the foreground mask images Image4_left and Image4_right of the power line-robot system.

[0067] Furthermore, S6 specifically includes:

[0068] Background replacement is performed based on texture fusion. The input image sequence consists of Image1_left and Image1_right, two new background images (Image3_left and Image3_right) from left and right perspectives, and two foreground mask images (Image4_left and Image4_right). The location of the foreground target of the power line-robot system in the wiring operation scene image is determined using the foreground mask image. Then, the pixels at the corresponding positions in the new background image are replaced with the pixels of the foreground target to complete the background replacement operation. The specific implementation method is as follows:

[0069] ;

[0070] in, for The pixel value of the location, For the new background image, Images of wiring operation scenarios. Foreground mask image, The initial composite image is output; the S6 operation is performed on the left and right viewpoints respectively to obtain the composite images Image5_left and Image5_right.

[0071] Furthermore, S7 specifically includes:

[0072] S71 performs morphological dilation and morphological erosion operations sequentially on the foreground mask images Image4_left and Image4_right to generate a Trimap three-color partition map. Within the gray unknown transition region of the Trimap three-color partition map, for each edge pixel, a continuous Alpha value between 0 and 1 is calculated using the Alpha Matting estimation algorithm. Given the black areas of a Trimap tri-color partition map white area The Alpha Matting estimation algorithm uses closed-form matting for any local pixel window in the image. Based on the linear model assumption, within the window Obtained from the linear affine function of the pixel RGB values:

[0073] ;

[0074] in, It is a three-dimensional linear coefficient vector. For scalar bias terms, Images of wiring operation scenarios. For window The index of any pixel contained within;

[0075] S72 solves for a clean foreground image using a foreground estimation algorithm. The foreground estimation algorithm uses a closed-form solution for foreground estimation, given... Under the premise of minimizing the synthesis error and the smoothing regularization term, a system of linear equations is constructed and solved. Its optimization objective is:

[0076] ;

[0077] in, For the new background image, subscript For any pixel position, This is the initial synthesized image. For gradient calculation; in the formula, the first term is the reconstruction error of image synthesis, and the second term is the smoothing regularization constraint. The regularization coefficient is used.

[0078] S73 for any pixel of the Trimap three-color partition map With formula Blend the foreground and new background to create a seamless, noise-free, smooth gradient edge. Optimize the resulting edge image;

[0079] S74 employs PCT-Net, a lightweight full-resolution image harmonization algorithm based on a dual-branch architecture, for color harmonization, optimizing the input edge image. With foreground mask image Synchronously downsample to a lower resolution to obtain a lower resolution input. and ; then and The input parameter network is modified based on the encoder-decoder U-shaped structure of iSSAM, with only the final output layer replaced. Convolutional layers; this parameter network outputs a low-resolution parameter map. ,in The set of real numbers, , The width and height of the low-resolution image. The number of color transformation parameters corresponding to a single pixel;

[0080] The S75 uses a bilinear interpolation algorithm to upsample the parameter map, transforming the low-resolution parameter map... Upsampled to the full-resolution input image With completely identical dimensions, a full-resolution parametric map is obtained. , , The width and height are the same as the dimensions of the full-resolution image, numerically equivalent to the wiring operation scene image. Width and height , same;

[0081] S76 For images composed of foreground mask images The defined background area, i.e. Directly retain the full-resolution input image The original pixel values ​​in the data, i.e. , Optimize images for color harmony;

[0082] S77 For images composed of foreground mask images The defined foreground area, i.e. Read the transformation parameters corresponding to this position. The harmonized pixel values ​​are calculated using the pixel-level color transformation PCT function: ;in The PCT transform function is an affine transformation type PCT function, and its transform form is as follows: ; where, linear transformation matrix With bias term They are respectively:

[0083] ;

[0084] ;

[0085] Each pixel corresponds to 12 parameters. Achieve linear mapping and color coupling adjustment between RGB channels. To achieve global brightness and hue shift;

[0086] Operations S71 to S77 are performed on the left and right viewpoints respectively, and finally the optimized images of the foreground and background from the left and right viewpoints, Image6_left and Image6_right, are obtained.

[0087] The beneficial effects of this invention are as follows:

[0088] 1. A simulation method for changing background images under dual-view consistency constraints is proposed.

[0089] 2. A shared yaw disturbance constraint is proposed to ensure that the binocular background view is consistent and the relative pose remains stable.

[0090] 3. A target-driven inverse sampling strategy is adopted to eliminate projection holes and improve the integrity of fisheye background generation.

[0091] 4. Construct a unified geometric representation of the spherical domain to achieve pixel-level precise geometric alignment between the foreground and background across different viewpoints.

[0092] 5. By integrating PyMatting and PCT-Net, the edges of the foreground and background are smoothed and the colors are coordinated and unified. Attached Figure Description

[0093] Figure 1 This is a flowchart of the method of the present invention. Detailed Implementation

[0094] This invention expands, recombines, and improves upon the new background generation and composite image optimization. While it shares many similarities with the original technology, it also incorporates numerous improvements, which are hereby explicitly stated. The invention will be further illustrated below with reference to specific embodiments.

[0095] Example 1:

[0096] This application provides a simulation method for changing the background of wiring images in a power grid robot based on dual-view consistency constraints. The flowchart of the method is as follows: Figure 1 As shown, the overall implementation path includes the following steps:

[0097] S1. Obtain the left and right view wiring operation scene image sequences Image1_left and Image1_right collected by the fisheye binocular vision system of the power grid robot, and construct a power grid robot wiring image sample library based on the image sequences;

[0098] S2 performs camera calibration on the image sequences Image1_left and Image1_right respectively, obtains the intrinsic parameters and distortion coefficients Data1_left and Data1_right of the left and right cameras respectively, calculates the consistency constraint relationship between the left and right dual views, and constructs the cross-view correlation mapping model Model1.

[0099] S3 unifies the background image Image2 to be replaced into a fisheye type image, then performs image denoising to obtain Image2_ready, and then maps Image2_ready and the image sequences Image1_left and Image1_right from the fisheye pixel plane to the unit sphere respectively;

[0100] S4 constructs a mapping relationship between Image2_ready and the target fisheye new background image through the unit spherical coordinate system, generates new backgrounds from the left and right perspectives, and obtains the visualized new background images Image3_left and Image3_right from the left and right perspectives.

[0101] S5 uses the U-Net segmentation network to perform semantic segmentation on the image sequences Image1_left and Image1_right, and extracts the foreground mask images Image4_left and Image4_right of the power line-robot system.

[0102] S6 The image sequences Image1_left and Image1_right, the new background images Image3_left and Image3_right from the left and right perspectives, and the foreground mask images Image4_left and Image4_right are respectively synthesized by combining the foreground images and new background images from the left and right perspectives to obtain the synthesized images Image5_left and Image5_right from the left and right perspectives.

[0103] S7 optimizes the foreground and background edges with jagged white edges and color and lighting inconsistencies in the synthesized images Image5_left and Image5_right, resulting in the final optimized images Image6_left and Image6_right.

[0104] Firstly, step S2 of the above scheme provides a binocular camera calibration method based on USM deep learning prediction and a cross-view correlation mapping modeling method based on shared rotation constraints. Camera calibration is performed on the fisheye binocular camera image sequences Image1_left and Image1_right. The traditional checkerboard calibration process is abandoned, and a unified spherical model (USM) parameter prediction method based on a single image is adopted. Using EfficientNetv2 as the backbone network, multi-task prediction of the field of view, nonlinear distortion parameters, attitude angle, etc., is performed. The intrinsic parameter matrices and distortion coefficients Data1_left and Data1_right of the left and right cameras are obtained analytically. Based on these two sets of parameters, a shared rotation and decoupled intrinsic parameter dual-view consistency constraint relationship is constructed. The background is modeled as an infinitely far spherical environment texture. Binocular view coordination is ensured through shared yaw disturbances. Finally, a cross-view correlation mapping model Model1 is established, providing accurate geometric constraints for subsequent background generation and preventing the destruction of binocular geometric relationships.

[0105] S2 includes the following:

[0106] S21. Based on a single-image wiring operation scenario, a Unified Spherical Model (USM) parameter prediction method is used to construct a multi-task prediction model with EfficientNetv2 as the backbone network. Given a wiring operation scenario image... The output of the prediction model network is represented as:

[0107] ;

[0108] in, , , and For continuous regression values, they correspond to the estimated values ​​of the field of view of the fisheye camera of the power grid robot, the estimated values ​​of the USM nonlinear distortion parameters, and the estimated values ​​of the pitch and roll angles, respectively. The pitch angle category prediction results are used to assist in the stable training of attitude estimation;

[0109] S22 constructs the focal length of the fisheye camera for the power grid robot through analytical calculation within the framework of the USM parameter prediction method. , Images of wiring operation scenarios width The field of view and USM nonlinear distortion parameters are used to represent:

[0110] ;

[0111] Image of wiring operation scene The principal point is taken as the image center. In the formula, , These are images of the wiring operation scene. The width and height are thus determined, allowing for assembly. Intrinsic parameter matrix:

[0112] ;

[0113] The output of the prediction model network in S21 After inverse normalization and scale recovery, the USM nonlinear distortion parameters in the true physical dimensions are obtained. Perform the construction evolution steps S21 to S22 on the image sequences Image1_left and Image1_right respectively to obtain the intrinsic parameter matrices of the left and right cameras. , and their respective distortion coefficients , This yields Data1_left and Data1_right;

[0114] S23 takes the pitch angle estimate output by the prediction model network in S21. Compared with the estimated roll angle After inverse normalization and scale recovery, the pitch angle in true physical dimensions is obtained. With roll angle The yaw angle is manually set according to the specific application scenario. ,according to - - Constructing rotation matrices in Euler angle order: Complete the attitude description of the camera in the world coordinate system;

[0115] S24 models the binocular vision system as a common pose perturbation under a rigid body structure and models the background image to be replaced as a spherical environment texture located at infinity.

[0116] S25 sets the left and right cameras to have basic yaw correction values ​​respectively. and This describes the fixed deviation of each installation orientation relative to a reference direction. For the background image to be replaced, this is used to obtain the first... A geometrically consistent background simulation sample is used, with a yaw disturbance shared by the left and right cameras applied. The rotation matrices of the left and right cameras are defined as follows:

[0117] ;

[0118] in, This represents the rotation matrix about the vertical axis. Representing the left or right camera, a cross-view correlation mapping model Model1 is constructed.

[0119] Secondly, step S3 of the above scheme provides a unified geometric representation method for preprocessing the background image to be replaced and mapping from fisheye to spherical surface. The background image to be replaced, Image2, is unified into a fisheye type image, and bilateral filtering is performed to preserve edges and denoise, eliminating compression noise and splicing artifacts to obtain Image2_ready; then, Image2_ready and the image sequences Image1_left and Image1_right are mapped from the fisheye pixel plane to the unit sphere, respectively, to achieve geometric alignment of the foreground and background in a unified spherical domain; providing high-quality and unified input data for spherical domain background generation.

[0120] S3 includes the following:

[0121] S31 sets the background image to be replaced. Image2 supports two image types: ERP panoramic image and standard pinhole image. The ERP panoramic image itself is captured by a fisheye camera and is a fisheye type image.

[0122] For the standard pinhole image type background image to be replaced, the isometric projection model commonly used in industrial fisheye cameras is used to convert it into a fisheye type image;

[0123] S32 uses bilateral filtering on the fisheye-type background image to be replaced. Apply edge-preserving noise reduction processing: for pixels at any position The denoising result is:

[0124] ;

[0125] in, For pixels The local neighborhood, For the neighborhood Any pixel within, and The smoothing intensity is controlled separately in the spatial domain and the grayscale domain. The normalization coefficient is... The denoised image is denoised as Image2_ready; the denoised ERP panoramic image is denoised as... A standard pinhole image, after fisheye conversion and denoising, is denoised as... ;

[0126] S33 for images Set its resolution to For any three-dimensional direction vector in the unit spherical coordinate system ,satisfy ,Will Convert to spherical angular coordinates:

[0127] ;

[0128] ;

[0129] in, for azimuth angle, for The pitch angle; where, To maintain consistency with the downward direction convention of the vertical axis in the ERP panoramic view coordinate system; then... Mapping from spherical angular coordinates to ERP 2D pixel texture coordinates:

[0130] ;

[0131] This establishes a mapping relationship between unit spherical angular coordinates and ERP two-dimensional pixel texture coordinates, based on which... Image pixels are desampled to a unit sphere;

[0132] S34 for images For any three-dimensional direction vector in the same unit spherical coordinate system , exist The normalized coordinates on the image plane are: Thus, a unit spherical coordinate system is constructed. The mapping relationship of two-dimensional pixel texture coordinates in an image, Regions whose projection points extend beyond the image boundaries are considered invalid regions; for valid regions, [the following is omitted as the text is incomplete and cannot be translated]. Image pixels are desampled to a unit sphere;

[0133] For wiring operation scenarios, the S35 inputs image pixel coordinates under a unified spherical model. Solve for the three-dimensional direction vector in the unit spherical coordinate system. :

[0134] ;

[0135] ;

[0136] ;

[0137] in, , , , Principal point of the image, For the focal length of the fisheye camera of the power grid robot, The nonlinear distortion parameters of the USM are used; the mapping from image pixel coordinates to unit spherical coordinates is denoted as the inverse mapping. Perform inverse mapping on the image sequences Image1_left and Image1_right respectively. The images of the wiring operation scene from both the left and right perspectives are projected from the fisheye pixel plane onto the unit sphere.

[0138] Thirdly, step S4 of the above scheme provides a method for collaborative generation of new backgrounds for binocular fisheye viewing based on viewpoint consistency constraints and spherical domain inverse sampling. In a unit spherical coordinate system, based on the cross-viewpoint correlation mapping model Model1 and shared yaw disturbance constraints, the spherical background is rotated and textured, adapting the independent intrinsic parameters of the left and right cameras to complete the generation of new backgrounds for both views. A target-driven inverse sampling strategy is adopted to avoid discrete holes generated by forward projection. The background texture in spherical coordinates is transformed back to the fisheye imaging plane through USM mapping, ultimately obtaining visual new background images Image3_left and Image3_right that conform to the geometric characteristics of binocular viewing and are consistent across viewpoints, providing geometrically matched background materials for foreground-background synthesis.

[0139] S4 includes the following:

[0140] S41 introduces a focal length scaling factor Intrinsic parameter matrix of fisheye camera for power grid robot Perform proportional scaling. Indicates the left or right camera, and obtains the adjusted intrinsic parameter matrix. :

[0141] ;

[0142] in, For construction The camera's focal length for Camera image principal point coordinates; Corresponding to the magnified focal length, Correspondingly, reduce the focal length. Then the original focal length remains unchanged;

[0143] S42 wiring operation scene image Resolution size is In order to obtain the same The target fisheye background images Image3_left and Image3_right are defined with the following pixel domains: ;

[0144] S43 Any pixel coordinate in the target fisheye background image Using the reverse mapping established in S35 Based on the parameters of the fisheye camera of the power grid robot Map it to a three-dimensional direction vector in the unit spherical coordinate system: ;in, Represents the unit sphere, and indicates the output vector. The unit line-of-sight direction vector in three-dimensional space;

[0145] S44 uses the rotation matrix constructed in S25. Apply a viewpoint rotation transformation, Transform the background texture orientation onto the unit sphere to obtain the rotated orientation vector: ;

[0146] S45, based on the source of the background image on the unit sphere, rotates the direction vector... Through different sampling mappings Convert to background image or The two-dimensional texture coordinates in the image; the background image on the unit sphere is derived from... hour, Calculated from S33; derived from hour, Calculated from S34;

[0147] S46 uses bilinear interpolation from the background image or The pixel values ​​are obtained to generate the new background image for the target fisheye.

[0148] S47 Different pixel coordinates of the target fisheye background image The pixel value at that location is given by the following formula:

[0149] ;

[0150] in, In the first Under a background simulation sample with consistent geometry from each viewpoint The camera's target fisheye background image; Image2_ready refers to the background image. or This results in two new background images, Image3_left and Image3_right, visualized from the left and right perspectives.

[0151] Fourthly, step S5 of the above scheme provides a method for extracting the foreground mask in a live-line operation scenario based on the U-Net segmentation network. The U-Net semantic segmentation network is used to perform pixel-level semantic segmentation on the image sequences Image1_left and Image1_right. Addressing the issues of large differences in multi-scale targets such as power lines, robotic arms, and utility poles, and radial distortion in fisheye images during live-line operation scenarios, U-Net skip connections are used to fuse shallow details and deep semantic features, preserving the continuous texture of thin wires and the complete outline of large targets. The network adapts to fisheye imaging characteristics, ensuring spatial consistency and edge accuracy of the segmentation results from both left and right perspectives, accurately extracting the foreground mask images Image4_left and Image4_right of the power line-robot system, providing precise geometric constraints for subsequent synthesis.

[0152] S5 includes the following:

[0153] S51 Constructs the power grid robot wiring image dataset Image1_dataset based on the image sequences Image1_left and Image1_right;

[0154] S52 Use image annotation software to annotate the power line-robot system region of a single image in the image dataset Image1_dataset, set it as a label to represent the category of the foreground region, save the annotation file in the same path as the image, and generate a JSON format annotation file containing the foreground region category;

[0155] S53 The labeled files are batch converted into binary mask format files required for training, and a semantic segmentation labeled dataset Labeling1_dataset containing original image-mask pairs is constructed.

[0156] S54. The semantic segmentation annotation dataset Labeling1_dataset is divided into training and testing sets. A U-Net network based on an encoder-decoder structure is constructed, and a focus loss function is built. and Dice coefficient loss Weighted Mixed Loss Function The expression is:

[0157] ;

[0158] ;

[0159] in, To balance the weighting coefficients, For the first The predicted probability value of each pixel. No. The actual label value of each pixel. Total number of pixels For smoothing terms;

[0160] S55 Iteratively trains the U-Net network using the training set and the hybrid loss function. When the Intersection over Union (IoU) index of the test set reaches a preset threshold, the optimal model weights are saved. The optimal model weights are loaded to perform inference on the image sequences Image1_left and Image1_right, and output the foreground mask images Image4_left and Image4_right of the power line-robot system.

[0161] Fifthly, step S6 of the above scheme provides a method for foreground and new background synthesis based on pixel-by-pixel replacement using a binary mask. This step takes the image sequences Image1_left and Image1_right, the new background images Image3_left and Image3_right from both left and right perspectives, and the foreground mask images Image4_left and Image4_right as input, and performs basic synthesis of the foreground and new background. A mask-driven texture fusion strategy is adopted, determining pixel affiliation based on the binary mask, accurately replacing the foreground target pixels with their corresponding positions in the new background, preserving the complete structure of the foreground and the global information of the background, and completing synchronous synthesis of the left and right perspectives according to the geometric relationship of the binocular viewpoints, resulting in the initial synthesized images Image5_left and Image5_right. This establishes the geometric structure and target position of the image, laying the foundation for subsequent optimization.

[0162] S6 includes the following:

[0163] Background replacement is performed based on texture fusion. The input image sequence consists of Image1_left and Image1_right, two new background images (Image3_left and Image3_right) from left and right perspectives, and two foreground mask images (Image4_left and Image4_right). The location of the foreground target of the power line-robot system in the wiring operation scene image is determined using the foreground mask image. Then, the pixels at the corresponding positions in the new background image are replaced with the pixels of the foreground target to complete the background replacement operation. The specific implementation method is as follows:

[0164] ;

[0165] in, for The pixel value of the location, For the new background image, Images of wiring operation scenarios. Foreground mask image, The initial composite image is output; the S6 operation is performed on the left and right viewpoints respectively to obtain the composite images Image5_left and Image5_right.

[0166] Sixthly, step S7 of the above scheme provides a method for optimizing synthesized images based on PyMatting and PCT-Net networks. To address the issues of jagged edges, white borders, and inconsistencies in foreground and background color and lighting in the initial synthesized images Image5_left and Image5_right, a two-stage optimization is implemented. Based on PyMatting, continuous alpha channels are estimated using Trimap and closed-form dematting to achieve a seamless and smooth transition between foreground and background edges, eliminating hard mask stitching marks. The lightweight PCT-Net color coordination network is used to predict color transformation parameters in low-resolution space and perform affine color transformation at full resolution, unifying the lighting, hue, and style of the foreground and background, ultimately obtaining visually realistic optimized images Image6_left and Image6_right.

[0167] S7 includes the following:

[0168] S71 performs morphological dilation and morphological erosion operations sequentially on the foreground mask images Image4_left and Image4_right to generate a Trimap three-color partition map. Within the gray unknown transition region of the Trimap three-color partition map, for each edge pixel, a continuous Alpha value between 0 and 1 is calculated using the Alpha Matting estimation algorithm. Given the black areas of a Trimap tri-color partition map white area The Alpha Matting estimation algorithm uses closed-form matting for any local pixel window in the image. Based on the linear model assumption, within the window Obtained from the linear affine function of the pixel RGB values:

[0169] ;

[0170] in, It is a three-dimensional linear coefficient vector. For scalar bias terms, Images of wiring operation scenarios. For window The index of any pixel contained within;

[0171] S72 solves for a clean foreground image using a foreground estimation algorithm. The foreground estimation algorithm uses a closed-form solution for foreground estimation, given... Under the premise of minimizing the synthesis error and the smoothing regularization term, a system of linear equations is constructed and solved. Its optimization objective is:

[0172] ;

[0173] in, For the new background image, subscript For any pixel position, This is the initial synthesized image. For gradient calculation; in the formula, the first term is the reconstruction error of image synthesis, and the second term is the smoothing regularization constraint. The regularization coefficient is used.

[0174] S73 for any pixel of the Trimap three-color partition map With formula Blend the foreground and new background to create a seamless, noise-free, smooth gradient edge. Optimize the resulting edge image;

[0175] S74 employs PCT-Net, a lightweight full-resolution image harmonization algorithm based on a dual-branch architecture, for color harmonization, optimizing the input edge image. With foreground mask image Synchronously downsample to a lower resolution to obtain a lower resolution input. and ; then and The input parameter network is modified based on the encoder-decoder U-shaped structure of iSSAM, with only the final output layer replaced. Convolutional layers; this parameter network outputs a low-resolution parameter map. ,in The set of real numbers, , The width and height of the low-resolution image. The number of color transformation parameters corresponding to a single pixel;

[0176] The S75 uses a bilinear interpolation algorithm to upsample the parameter map, transforming the low-resolution parameter map... Upsampled to the full-resolution input image With completely identical dimensions, a full-resolution parametric map is obtained. , , The width and height are the same as the dimensions of the full-resolution image, numerically equivalent to the wiring operation scene image. Width and height , same;

[0177] S76 For images composed of foreground mask images The defined background area, i.e. Directly retain the full-resolution input image The original pixel values ​​in the data, i.e. , Optimize images for color harmony;

[0178] S77 For images composed of foreground mask images The defined foreground area, i.e. Read the transformation parameters corresponding to this position. The harmonized pixel values ​​are calculated using the pixel-level color transformation PCT function: ;in The PCT transform function is an affine transformation type PCT function, and its transform form is as follows: ; where, linear transformation matrix With bias term They are respectively:

[0179] ;

[0180] ;

[0181] Each pixel corresponds to 12 parameters. Achieve linear mapping and color coupling adjustment between RGB channels. To achieve global brightness and hue shift;

[0182] Operations S71 to S77 are performed on the left and right viewpoints respectively, and finally the optimized images of the foreground and background from the left and right viewpoints, Image6_left and Image6_right, are obtained.

[0183] This application presents a simulation method for background replacement in power grid robot wiring images based on dual-view consistency constraints. This method combines view consistency constraints, associative new background generation, semantic segmentation, foreground and new background synthesis, and optimization techniques to construct an integrated framework for image background replacement in power grid robot wiring operation scenarios. For wiring operation scene images acquired by the power grid robot's fisheye binocular vision system, firstly, a view consistency constraint relationship is constructed based on camera imaging parameters to establish a cross-view correlation mapping model. Next, the background image to be replaced undergoes fisheye view transformation and is uniformly mapped to a unit spherical coordinate system for geometric alignment. Then, combined with view constraints, a new background is generated in the spherical domain and reprojected back onto the fisheye imaging plane to obtain new backgrounds from the left and right perspectives that match the binocular pose and have consistent view relationships. Then, the original binocular fisheye image is semantically segmented using a U-Net network to extract the power line-robot system foreground mask. Finally, based on the mask, the foreground and new background of the left and right camera images are synthesized. Simultaneously, edge smoothing and color harmonization optimization processes are performed to address issues such as jagged white edges and inconsistencies in foreground and background color and illumination in the synthesized image. Compared with traditional single-view independent background replacement methods, this invention can strictly guarantee the geometric consistency of the binocular background, effectively realize dual-view collaborative background replacement, construct a high-fidelity and diverse power grid operation simulation dataset, and significantly improve the training effect and scene generalization ability of the visual model.

Claims

1. A simulation method for background replacement in wiring images of power grid robots based on dual-view consistency constraints, characterized in that, The method includes the following steps: S1. Obtain the left and right view wiring operation scene image sequences Image1_left and Image1_right collected by the fisheye binocular vision system of the power grid robot, and construct a power grid robot wiring image sample library based on the image sequences; S2 performs camera calibration on the image sequences Image1_left and Image1_right respectively, obtains the intrinsic parameters and distortion coefficients Data1_left and Data1_right of the left and right cameras respectively, calculates the consistency constraint relationship between the left and right dual views, and constructs the cross-view correlation mapping model Model1. S3 unifies the background image Image2 to be replaced into a fisheye type image, then performs image denoising to obtain Image2_ready, and then maps Image2_ready and the image sequences Image1_left and Image1_right from the fisheye pixel plane to the unit sphere respectively; S4 constructs a mapping relationship between Image2_ready and the target fisheye new background image through the unit spherical coordinate system, generates new backgrounds from the left and right perspectives, and obtains the visualized new background images Image3_left and Image3_right from the left and right perspectives. S5 uses the U-Net segmentation network to perform semantic segmentation on the image sequences Image1_left and Image1_right, and extracts the foreground mask images Image4_left and Image4_right of the power line-robot system. S6 The image sequences Image1_left and Image1_right, the new background images Image3_left and Image3_right from the left and right perspectives, and the foreground mask images Image4_left and Image4_right are respectively synthesized by combining the foreground images and new background images from the left and right perspectives to obtain the synthesized images Image5_left and Image5_right from the left and right perspectives. S7 optimizes the foreground and background edges with jagged white edges and color and lighting inconsistencies in the synthesized images Image5_left and Image5_right, resulting in the final optimized images Image6_left and Image6_right.

2. The simulation method for background replacement of wiring images for power grid robots based on dual-view consistency constraints according to claim 1, characterized in that, Specifically, S2 is: S21. Based on a single-image wiring operation scenario, a Unified Spherical Model (USM) parameter prediction method is used to construct a multi-task prediction model with EfficientNetv2 as the backbone network. Given a wiring operation scenario image... The output of the prediction model network is represented as: ; in, , , and For continuous regression values, they correspond to the estimated values ​​of the field of view of the fisheye camera of the power grid robot, the estimated values ​​of the USM nonlinear distortion parameters, and the estimated values ​​of the pitch and roll angles, respectively. The pitch angle category prediction results are used to assist in the stable training of attitude estimation; S22 constructs the focal length of the fisheye camera for the power grid robot through analytical calculation within the framework of the USM parameter prediction method. , Images of wiring operation scenarios width The field of view and USM nonlinear distortion parameters are used to represent: ; Image of wiring operation scene The principal point is taken as the image center. In the formula, , These are images of the wiring operation scene. The width and height are thus determined, allowing for assembly. Intrinsic parameter matrix: ; The output of the prediction model network in S21 After inverse normalization and scale recovery, the USM nonlinear distortion parameters in the true physical dimensions are obtained. Perform the construction evolution steps S21 to S22 on the image sequences Image1_left and Image1_right respectively to obtain the intrinsic parameter matrices of the left and right cameras. , and their respective distortion coefficients , This yields Data1_left and Data1_right; S23 takes the pitch angle estimate output by the prediction model network in S21. Compared with the estimated roll angle After inverse normalization and scale recovery, the pitch angle in true physical dimensions is obtained. With roll angle The yaw angle is manually set according to the specific application scenario. ,according to - - Constructing rotation matrices in Euler angle order: Complete the attitude description of the camera in the world coordinate system; S24 models the binocular vision system as a common pose perturbation under a rigid body structure and models the background image to be replaced as a spherical environment texture located at infinity. S25 sets the left and right cameras to have basic yaw correction values ​​respectively. and This describes the fixed deviation of each installation orientation relative to a reference direction. For the background image to be replaced, this is used to obtain the first... A geometrically consistent background simulation sample is used, with a yaw disturbance shared by the left and right cameras applied. The rotation matrices of the left and right cameras are defined as follows: ; in, This represents the rotation matrix about the vertical axis. Representing the left or right camera, a cross-view correlation mapping model Model1 is constructed.

3. The simulation method for background replacement of wiring images for power grid robots based on dual-view consistency constraints according to claim 2, characterized in that, Specifically, S3 is: S31 sets the background image to be replaced. Image2 supports two image types: ERP panoramic image and standard pinhole image. The ERP panoramic image itself is captured by a fisheye camera and is a fisheye type image. For the standard pinhole image type background image to be replaced, the isometric projection model commonly used in industrial fisheye cameras is used to convert it into a fisheye type image; S32 uses bilateral filtering on the fisheye-type background image to be replaced. Apply edge-preserving noise reduction processing: for pixels at any position The denoising result is: ; in, For pixels The local neighborhood, For the neighborhood Any pixel within, and The smoothing intensity is controlled separately in the spatial domain and the grayscale domain. The normalization coefficient is... The denoised image is denoised as Image2_ready; the denoised ERP panoramic image is denoised as... A standard pinhole image, after fisheye conversion and denoising, is denoised as... ; S33 for images Set its resolution to For any three-dimensional direction vector in the unit spherical coordinate system ,satisfy ,Will Convert to spherical angular coordinates: ; ; in, for azimuth angle, for The pitch angle; where, To maintain consistency with the downward direction convention of the vertical axis in the ERP panoramic view coordinate system; then... Mapping from spherical angular coordinates to ERP 2D pixel texture coordinates: ; This establishes a mapping relationship between unit spherical angular coordinates and ERP two-dimensional pixel texture coordinates, based on which... Image pixels are desampled to a unit sphere; S34 for images For any three-dimensional direction vector in the same unit spherical coordinate system , exist The normalized coordinates on the image plane are: Thus, a unit spherical coordinate system is constructed. The mapping relationship of two-dimensional pixel texture coordinates in an image, Regions whose projection points extend beyond the image boundaries are considered invalid regions; for valid regions, [the following is omitted as the text is incomplete and cannot be translated]. Image pixels are desampled to a unit sphere; For wiring operation scenarios, the S35 inputs image pixel coordinates under a unified spherical model. Solve for the three-dimensional direction vector in the unit spherical coordinate system. : ; ; ; in, , , , Principal point of the image, For the focal length of the fisheye camera of the power grid robot, The nonlinear distortion parameters of the USM are used; the mapping from image pixel coordinates to unit spherical coordinates is denoted as the inverse mapping. Perform inverse mapping on the image sequences Image1_left and Image1_right respectively. The images of the wiring operation scene from both the left and right perspectives are projected from the fisheye pixel plane onto the unit sphere.

4. The simulation method for changing the background of wiring images for power grid robots based on dual-view consistency constraints according to claim 3, characterized in that, Specifically, S4 is: S41 introduces a focal length scaling factor Intrinsic parameter matrix of fisheye camera for power grid robot Perform proportional scaling. Indicates the left or right camera, and obtains the adjusted intrinsic parameter matrix. : ; in, For construction The camera's focal length for Camera image principal point coordinates; Corresponding to the magnified focal length, Correspondingly, reduce the focal length. Then the original focal length remains unchanged; S42 wiring operation scene image Resolution size is In order to obtain the same The target fisheye background images Image3_left and Image3_right are defined with the following pixel domains: ; S43 Any pixel coordinate in the target fisheye background image Using the reverse mapping established in S35 Based on the parameters of the fisheye camera of the power grid robot Map it to a three-dimensional direction vector in the unit spherical coordinate system: ;in, Represents the unit sphere, and indicates the output vector. The unit line-of-sight direction vector in three-dimensional space; S44 uses the rotation matrix constructed in S25. Apply a viewpoint rotation transformation, Transform the background texture orientation onto the unit sphere to obtain the rotated orientation vector: ; S45, based on the source of the background image on the unit sphere, rotates the direction vector... Through different sampling mappings Convert to background image or The two-dimensional texture coordinates in the image; the background image on the unit sphere is derived from... hour, Calculated from S33; derived from hour, Calculated from S34; S46 uses bilinear interpolation from the background image or The pixel values ​​are obtained to generate the new background image for the target fisheye. S47 Different pixel coordinates of the target fisheye background image The pixel value at that location is given by the following formula: ; in, In the first Under a background simulation sample with consistent geometry from each viewpoint The camera's target fisheye background image; Image2_ready refers to the background image. or This results in two new background images, Image3_left and Image3_right, visualized from the left and right perspectives.

5. The simulation method for changing the background of wiring images for power grid robots based on dual-view consistency constraints according to claim 4, characterized in that, Specifically, S5 is: S51 Constructs the power grid robot wiring image dataset Image1_dataset based on the image sequences Image1_left and Image1_right; S52 Use image annotation software to annotate the power line-robot system region of a single image in the image dataset Image1_dataset, set it as a label to represent the category of the foreground region, save the annotation file in the same path as the image, and generate a JSON format annotation file containing the foreground region category; S53 The labeled files are batch converted into binary mask format files required for training, and a semantic segmentation labeled dataset Labeling1_dataset containing original image-mask pairs is constructed. S54. The semantic segmentation annotation dataset Labeling1_dataset is divided into training and testing sets. A U-Net network based on an encoder-decoder structure is constructed, and a focus loss function is built. and Dice coefficient loss Weighted Mixed Loss Function The expression is: ; ; in, To balance the weighting coefficients, For the first The predicted probability value of each pixel. No. The actual label value of each pixel. Total number of pixels For smoothing terms; S55 Iteratively trains the U-Net network using the training set and the hybrid loss function. When the Intersection over Union (IoU) index of the test set reaches a preset threshold, the optimal model weights are saved. The optimal model weights are loaded to perform inference on the image sequences Image1_left and Image1_right, and output the foreground mask images Image4_left and Image4_right of the power line-robot system.

6. The simulation method for changing the background of wiring images for power grid robots based on dual-view consistency constraints according to claim 5, characterized in that, Specifically, S6 is: Background replacement is performed based on texture fusion. The input image sequence consists of Image1_left and Image1_right, two new background images (Image3_left and Image3_right) from left and right perspectives, and two foreground mask images (Image4_left and Image4_right). The location of the foreground target of the power line-robot system in the wiring operation scene image is determined using the foreground mask image. Then, the pixels at the corresponding positions in the new background image are replaced with the pixels of the foreground target to complete the background replacement operation. The specific implementation method is as follows: ; in, for The pixel value of the location, For the new background image, Images of wiring operation scenarios. Foreground mask image, The initial composite image is output; the S6 operation is performed on the left and right viewpoints respectively to obtain the composite images Image5_left and Image5_right.

7. The simulation method for changing the background of a power grid robot wiring image based on dual-view consistency constraints according to claim 6, characterized in that, Specifically, S7 is: S71 performs morphological dilation and morphological erosion operations sequentially on the foreground mask images Image4_left and Image4_right to generate a Trimap three-color partition map. Within the gray unknown transition region of the Trimap three-color partition map, for each edge pixel, a continuous Alpha value between 0 and 1 is calculated using the Alpha Matting estimation algorithm. Given the black areas of a Trimap tri-color partition map white area The Alpha Matting estimation algorithm uses closed-form matting for any local pixel window in the image. Based on the linear model assumption, within the window Obtained from the linear affine function of the pixel RGB values: ; in, It is a three-dimensional linear coefficient vector. For scalar bias terms, Images of wiring operation scenarios. For window The index of any pixel contained within; S72 solves for a clean foreground image using a foreground estimation algorithm. The foreground estimation algorithm uses a closed-form solution for foreground estimation, given... Under the premise of minimizing the synthesis error and the smoothing regularization term, a system of linear equations is constructed and solved. Its optimization objective is: ; in, For the new background image, subscript For any pixel position, This is the initial synthesized image. For gradient calculation; in the formula, the first term is the reconstruction error of image synthesis, and the second term is the smoothing regularization constraint. The regularization coefficient is used. S73 for any pixel of the Trimap three-color partition map With formula Blend the foreground and new background to create a seamless, noise-free, smooth gradient edge. Optimize the resulting edge image; S74 employs PCT-Net, a lightweight full-resolution image harmonization algorithm based on a dual-branch architecture, for color harmonization, optimizing the input edge image. With foreground mask image Synchronously downsample to a lower resolution to obtain a lower resolution input. and ; then and The input parameter network is modified based on the encoder-decoder U-shaped structure of iSSAM, with only the final output layer replaced. Convolutional layers; this parameter network outputs a low-resolution parameter map. ,in The set of real numbers, , The width and height of the low-resolution image. The number of color transformation parameters corresponding to a single pixel; The S75 uses a bilinear interpolation algorithm to upsample the parameter map, transforming the low-resolution parameter map... Upsampled to the full-resolution input image With completely identical dimensions, a full-resolution parametric map is obtained. , , The width and height are the same as the dimensions of the full-resolution image, numerically equivalent to the wiring operation scene image. Width and height , same; S76 For images composed of foreground mask images The defined background area, i.e. Directly retain the full-resolution input image The original pixel values ​​in the data, i.e. , Optimize images for color harmony; S77 For images composed of foreground mask images The defined foreground area, i.e. Read the transformation parameters corresponding to this position. The harmonized pixel values ​​are calculated using the pixel-level color transformation PCT function: ;in The PCT transform function is an affine transformation type PCT function, and its transform form is as follows: ; where, linear transformation matrix With bias term They are respectively: ; ; Each pixel corresponds to 12 parameters. Achieve linear mapping and color coupling adjustment between RGB channels. To achieve global brightness and hue shift; Operations S71 to S77 are performed on the left and right viewpoints respectively, and finally the optimized images of the foreground and background from the left and right viewpoints, Image6_left and Image6_right, are obtained.