Autonomous solar power generation installation using artificial intelligence

The solar panel handling system addresses installation challenges by employing an arm assembly tool with machine learning to adapt to environmental conditions, ensuring precise and efficient panel alignment and reducing costs.

JP2025529762APending Publication Date: 2025-09-09ジ·エーイーエス·コーポレーション
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025507644
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-08-11
Filing Date
2023-08-11
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

The installation of solar panels on a mounting structure is challenging due to their fragility and large size, and traditional computer vision techniques are hindered by glare and lighting issues, leading to inefficiencies and increased costs.

Method used

A solar panel handling system using an arm assembly tool with a suction cup, linear guide assembly, and force-torque transducer, combined with machine learning techniques to overcome environmental inconsistencies and ensure precise installation.

Benefits of technology

The system enables efficient and reliable installation of solar panels by utilizing machine learning to adapt to environmental conditions, ensuring accurate alignment and reducing installation costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025529762000001_ABST
    Figure 2025529762000001_ABST
Patent Text Reader

Abstract

A system and method for installing solar panels is provided. The method includes acquiring images of the solar panels and an installation structure during installation. The method further includes preprocessing the images by correcting for camera inherent characteristics or distortions, rectifying the images, and / or determining depth information. The method further includes detecting the solar panels by inputting the images into a neural network. The method further includes first post-processing to calculate a first panel pose based on an output of the neural network. The method further includes generating a control signal for operating a robotic controller for installing the solar panels based on the first panel pose. In some embodiments, the method further includes performing a homography transformation to obtain a second panel pose based on the first panel pose and a visual pattern or visual reference on the solar panel, and further generating the control signal based on the second panel pose.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] Related application data This application is based on and claims priority under U.S. Provisional Patent Application No. 63 / 397,125, filed August 11, 2022, to that provisional patent under 35 U.S.C. § 119, the entire contents of which are incorporated herein by reference.

[0002] The present disclosure relates generally to solar panel handling systems, and more particularly to systems and methods for installing solar panels on a mounting structure. [Background technology]

[0003] In the following description, reference is made to certain structures and / or methods. However, the following reference should not be construed as an admission that these structures and / or methods constitute prior art. Applicant expressly reserves the right to demonstrate that such structures and / or methods do not qualify as prior art against the present invention.

[0004] Solar array installation typically involves securing solar panels to a mounting structure. This base support not only provides mounting points for the individual solar panels but also assists in routing the electrical system and, if applicable, any mechanical components. Due to the fragility and large size of solar panels, the process of securing solar panels to a mounting structure presents unique challenges. For example, in many instances, solar panels in a solar array are mounted on a rotatable structure. This structure rotates the solar panels around an axis, allowing the array to track the sun. In such instances, it is difficult to ensure that all solar panels in the array are flush and level with the axis of the rotatable structure. Furthermore, the installation cost of a solar array can account for a significant portion of the overall cost of building a solar array. Therefore, there is a need for a more efficient and reliable solar panel handling system for installing solar panels within a solar array. In ideal environments, traditional computer vision techniques can be utilized. However, glare, overexposure, or underexposure can adversely affect object detection algorithms. [Prior art documents] [Non-patent literature]

[0005] [Non-Patent Document 1] J. Spencer et al., "The monocular depth estimation challenge," Proceedings of the IEEE / CVF Winter Conference on Applications of Computer Vision, (pp. 623-632), (2023) [Non-patent document 2] H. Hirschmuller, "Stereo Processing by Semiglobal Matching and Mutual Information," IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 30, no. 2, pp. 328-341, February 2008, doi: 10.1109 / TPAMI.2007.1166 [Non-patent document 3] S. Patil et al., “A Comparative Evaluation of SGM Variants (including a New Variant, tMGM) for Dense Stereo Matching,” ArXiv, abs / 1911.09800 (2019) [Non-patent document 4] N. O'Mahony et al., "Deep learning vs. traditional computer vision," Advances in Computer Vision: Proceedings of the 2019 Computer Vision Conference (CVC), Springer Nature Switzerland AG, Volume 11 (pp. 128-144) (2020) Summary of the Invention

[0006] Accordingly, the present invention is directed to a solar panel handling system that substantially obviates one or more of the problems resulting from limitations and drawbacks of the related art.

[0007] The solar panel handling system disclosed herein facilitates the installation of solar panels of a solar array onto an existing installation structure, such as a torque tube. By combining tools for handling solar panels with components that allow for coupling of the solar panels to the solar panel support structure, solar panel installation can be made more efficient and reliable. Some embodiments utilize machine learning techniques to overcome environmental inconsistencies. The system can learn from examples with glare and lighting issues and have the ability to generalize to new data during inference.

[0008] Additional features and advantages of the invention will be set forth in the description which follows, and in part will be obvious from the description, or may be learned by the practice of the invention. The objectives and other advantages of the invention may be realized and attained by the structure particularly pointed out in the detailed description and claims hereof, as well as the appended drawings. [Means for solving the problem]

[0009] To achieve these and other advantages and in accordance with the purposes of the present invention, as embodied and broadly described, a system for installing solar panels may include an arm assembly tool end including a frame and a suction cup coupled to the frame, a linear guide assembly coupled to the arm assembly tool end, the linear guide assembly including a linearly movable clamping tool including an engagement member configured to engage a clamping assembly slidably coupled to the installation structure, a force-torque transducer configured to move the clamping tool along the installation structure, and a junction box coupled to the frame including a controller and power source configured to control the force-torque transducer and the suction cup.

[0010] In another aspect, a method of installing a solar panel may include engaging an arm assembly tool end with a solar panel, the arm assembly tool end comprising a frame and a suction cup coupled to the frame; positioning the solar panel relative to a mounting structure to which a clamping assembly is slidably coupled; engaging a linear guide assembly coupled to the arm assembly tool end with a clamping assembly, the linear guide assembly comprising a linearly movable clamping tool comprising an engagement member configured to engage the clamping assembly and a force-torque converter configured to move the clamping tool along the mounting structure; and actuating the force-torque converter to move the clamping assembly along the mounting structure to engage a side of the solar panel, thereby securing the solar panel relative to the mounting structure.

[0011] In another aspect, a method for training a neural network for autonomous solar power installation is provided according to some embodiments. The method includes acquiring one or more images during installation. The one or more images include images of one or more solar panels and an installation structure. The method further includes preprocessing the one or more images, including one or more of correcting for camera inherent characteristics or distortions, rectifying the images, and determining depth information. The method further includes detecting the one or more solar panels by inputting the one or more images into one or more neural networks trained to detect solar panels. The method further includes first post-processing for calculating a first panel pose based on the output of the one or more neural networks. The method further includes generating control signals for operating a robotic controller to install the one or more solar panels based on the first panel pose. In some embodiments, the method further includes second post-processing including one or more homography transformations for obtaining second panel poses of the one or more solar panels based on the first panel pose. The second post-processing corrects or corrects inaccuracies in the first panel pose based on visual patterns or visual references on the solar panels. These control signals for operating the robotic controller to place the one or more solar panels are further generated based on the pose of the second panel.

[0012] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory and are intended to provide further explanation of the invention as claimed.

[0013] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate the invention and, together with the description, further serve to explain the principles of the invention and to enable one skilled in the relevant art to make and use the invention. The illustrative embodiments can be best understood from the following detailed description when read in conjunction with the accompanying drawings. It is emphasized that, according to common practice, the various features of the drawings are not to scale. To the contrary, dimensions of the various features have been arbitrarily expanded or reduced for clarity. The drawings include the following figures: [Brief explanation of the drawings]

[0014] [Figure 1] FIG. 1 is a perspective view of a solar panel handling system with a container of solar panels according to one embodiment of the present disclosure. [Figure 2A] FIG. 2 is a plan view of the solar panel handling system and solar panel container of FIG. 1. [Figure 2B] FIG. 2 is a front view of the solar panel handling system and solar panel container of FIG. 1. [Figure 2C] FIG. 2 is a side view of the solar panel handling system and solar panel container of FIG. 1. [Figure 3A] FIG. 1 is a plan view of a solar panel handling system coupled to a single solar panel, according to one embodiment of the present disclosure. [Figure 3B] FIG. 1 illustrates a front view of a solar panel handling system coupled to a single solar panel, according to one embodiment of the present disclosure. [Figure 3C] FIG. 1 is a side view of a solar panel handling system coupled to a single solar panel according to one embodiment of the present disclosure. [Figure 4A] FIG. 1 is a perspective view of a solar panel handling system according to one embodiment of the present disclosure. [Figure 4B] FIG. 1 is a perspective view of a solar panel handling system according to one embodiment of the present disclosure. [Figure 5A]FIG. 1 is a plan view of a solar panel handling system according to one embodiment of the present disclosure. [Figure 5B] FIG. 1 is a front view of a solar panel handling system according to one embodiment of the present disclosure. [Figure 5C] FIG. 1 illustrates a side view of a clamping tool of a solar panel handling system in a stowed position according to one embodiment of the present disclosure. [Figure 5D] FIG. 10 is a side view of a clamping tool in an extended or advanced position according to one embodiment of the present disclosure. [Figure 6A] FIG. 1 is a perspective view of a clamping tool of a solar panel handling system in engagement with a clamping assembly coupled to a mounting structure according to one embodiment of the present disclosure. [Figure 6B] FIG. 1 is a perspective view of a clamping tool of a solar panel handling system in engagement with a clamping assembly coupled to a mounting structure according to one embodiment of the present disclosure. [Figure 7A] FIG. 10 is a plan view of a clamping tool of a solar panel handling system in engagement with a clamping assembly coupled to a mounting structure according to one embodiment of the present disclosure. [Figure 7B] FIG. 1 is a front view of a clamping tool of a solar panel handling system engaged with a clamping assembly coupled to a mounting structure according to one embodiment of the present disclosure. [Figure 7C] FIG. 1 is a side view of a clamping tool of a solar panel handling system in engagement with a clamping assembly coupled to a mounting structure according to one embodiment of the present disclosure. [Figure 7D] FIG. 10 is a rear view of a clamping tool of a solar panel handling system engaged with a clamping assembly coupled to a mounting structure according to one embodiment of the present disclosure. [Figure 8] FIG. 1 is a schematic diagram of an overhead view of a solar panel handling system during the process of installing a solar panel, according to one embodiment of the present disclosure. [Figure 9]FIG. 1 illustrates a solar panel handling system including an assembly tool coupled to an assembly mobile robot using a robotic arm. [Figure 10] FIG. 1 illustrates a solar panel handling system with two robotic arms, with two assembly tools coupled to an assembly mobile robot using their respective robotic arms. [Figure 11A] FIG. 1 illustrates a process for installing solar panels. [Figure 11B] FIG. 1 illustrates a process for installing solar panels. [Figure 11C] FIG. 1 illustrates a process for installing solar panels. [Figure 12A] FIG. 1 is a diagram showing the configuration of a mobile robot system including two modular vehicles and a land vehicle with two robotic arms. [Figure 12B] FIG. 1 is a diagram showing the configuration of a mobile robot system including two modular vehicles and a land vehicle with two robotic arms. [Figure 13] FIG. 1 is a schematic diagram of placement achieved using computer vision registration. [Figure 14] FIG. 1 is a schematic diagram of an arrangement in which a modular vehicle is replaced with a new modular vehicle having additional supplemental solar panels. [Figure 15] FIG. 1 is a detailed diagram of an example configuration of a system for installing solar panels according to an embodiment of the present disclosure. [Figure 16] FIG. 1 is a detailed diagram of an example configuration of a system for installing solar panels according to an embodiment of the present disclosure. [Figure 17] FIG. 1 is a detailed diagram of an example configuration of a system for installing solar panels according to an embodiment of the present disclosure. [Figure 18] FIG. 1 is a detailed diagram of an example configuration of a system for installing solar panels according to an embodiment of the present disclosure. [Figure 19] FIG. 1 is a detailed diagram of an example configuration of a system for installing solar panels according to an embodiment of the present disclosure. [Figure 20] FIG. 1 is a detailed diagram of an example configuration of a system for installing solar panels according to an embodiment of the present disclosure. [Figure 21A] FIG. 1 is a detailed diagram of an example configuration of a system for installing solar panels according to an embodiment of the present disclosure. [Figure 21B] FIG. 1 is a detailed diagram of an example configuration of a system for installing solar panels according to an embodiment of the present disclosure. [Figure 22] FIG. 1 is a detailed diagram of an example configuration of a system for installing solar panels according to an embodiment of the present disclosure. [Figure 23] FIG. 1 is a detailed diagram of an example configuration of a system for installing solar panels according to an embodiment of the present disclosure. [Figure 24] FIG. 1 is a detailed diagram of an example configuration of a system for installing solar panels according to an embodiment of the present disclosure. [Figure 25] FIG. 1 is a detailed diagram of an example configuration of a system for installing solar panels according to an embodiment of the present disclosure. [Figure 26] FIG. 1 is a detailed diagram of an example configuration of a system for installing solar panels according to an embodiment of the present disclosure. [Figure 27] FIG. 1 is a detailed diagram of an example configuration of a system for installing solar panels according to an embodiment of the present disclosure. [Figure 28] FIG. 1 is a detailed diagram of an example configuration of a system for installing solar panels according to an embodiment of the present disclosure. [Figure 29] FIG. 1 is a detailed diagram of an example configuration of a system for installing solar panels according to an embodiment of the present disclosure. [Figure 30] FIG. 1 is a detailed diagram of an example configuration of a system for installing solar panels according to an embodiment of the present disclosure. [Figure 31] FIG. 1 is a detailed diagram of an example configuration of a system for installing solar panels according to an embodiment of the present disclosure. [Figure 32] FIG. 1 is a detailed diagram of an example configuration of a system for installing solar panels according to an embodiment of the present disclosure. [Figure 33]FIG. 1 is a detailed diagram of an example configuration of a system for installing solar panels according to an embodiment of the present disclosure. [Figure 34] FIG. 1 is a detailed diagram of an example configuration of a system for installing solar panels according to an embodiment of the present disclosure. [Figure 35A] FIG. 1 is a block diagram of an example image processing pipeline in accordance with some embodiments. [Figure 35B] FIG. 1 illustrates an example captured corrected image according to some embodiments. [Figure 35C] FIG. 35C illustrates the output of an example of neural network image segmentation of the acquired corrected image shown in FIG. 35B, according to some embodiments. [Figure 35D] FIG. 1 illustrates an example panel corner detection according to some embodiments. [Figure 36] 1A-1C illustrate examples of images of roads under different lighting conditions and segmentation masks for those images, according to some embodiments. [Figure 37A] FIG. 1 illustrates an example of a captured image including a solar panel and a torque tube, according to some embodiments. [Figure 37B] FIG. 37B illustrates an example of an annotated image of the captured image shown in FIG. 37A, according to some embodiments. [Figure 38A] FIG. 10 is a diagram illustrating an example of image classification. [Figure 38B] FIG. 38B shows an example of object localization for the image shown in FIG. 38A. [Figure 38C] FIG. 1 illustrates an example of semantic segmentation according to some embodiments. [Figure 39] FIG. 1 illustrates an example of instance segmentation of a solar panel according to some embodiments. [Figure 40] FIG. 1 illustrates an example image handling system according to some embodiments. [Figure 41] FIG. 1 illustrates a trailer system according to some embodiments. [Figure 42A]FIG. 10 illustrates a histogram of the norm of the pose error of a neural network when using coarse position, according to some implementations. [Figure 42B] FIG. 10 illustrates a histogram of the norm of the pose error of a neural network without coarse position, according to some implementations. [Figure 43A] FIG. 1 illustrates an example of the approximate location of solar panels using the B Mask R-CNN model, according to some embodiments. [Figure 43B] FIG. 1 illustrates an example of the approximate location of solar panels using the B Mask R-CNN model, according to some embodiments. [Figure 44] FIG. 44 illustrates a system 4400 for solar panel installation according to some embodiments. [Figure 45A] FIG. 1 illustrates a vision system for tracking the position of a trailer. [Figure 45B] FIG. 1 illustrates an expanded view of a vision system in accordance with some embodiments. [Figure 46A] FIG. 1 illustrates a vision system for module picking according to some embodiments. [Figure 46B] FIG. 1 illustrates an expanded view of a vision system in accordance with some embodiments. [Figure 47A] FIG. 47 illustrates a system 4700 for distance measurement at a modular angle according to some embodiments. [Figure 47B] FIG. 47 illustrates a system 4700 for distance measurement at a modular angle according to some embodiments. [Figure 47C] FIG. 47 illustrates a system 4700 for distance measurement at a modular angle according to some embodiments. [Figure 48A] FIG. 1 illustrates a system for generating laser lines to detect the position of the tube and clamp, according to some embodiments. [Figure 48B] FIG. 48B is an enlarged view of the laser line generating system shown in FIG. 48A. [Figure 48C]FIG. 10 illustrates laser line generation according to some embodiments (horizontal lines detect clamps, vertical lines detect tubes). [Figure 49A] FIG. 49 illustrates a vision system 4900 for estimating the position of the tube and clamp. [Figure 49B] FIG. 1 illustrates an expanded view of a vision system in accordance with some embodiments. [Figure 50A] 1 is a flow diagram of an autonomous solar power plant installation according to some embodiments. [Figure 50B] 1 is a flow diagram of a method for training a neural network for autonomous solar power installations according to some embodiments. [Figure 51] FIG. 1 illustrates an example application of the Hough transform to identify corners of a solar panel, according to some embodiments. [Figure 52] FIG. 1 illustrates an example application of segmentation according to some embodiments. [Figure 53] FIG. 1 illustrates an example application of a corner detection algorithm according to some embodiments. [Figure 54] FIG. 10 illustrates grid intersections that can be detected using a two-pass homography algorithm, according to some embodiments. [Figure 55] FIG. 1 is a schematic diagram of region of interest segmentation according to some embodiments. [Figure 56] FIG. 1 illustrates multi-data channel performance of a LiDAR used in accordance with some embodiments. [Figure 57] 1 is a schematic diagram of an example computer vision / artificial intelligence (AI) architecture for estimating six degrees of freedom (6DoF) pose of a solar panel, according to some embodiments. [Figure 58] FIG. 1 is a schematic diagram of another example computer vision / AI pipeline for estimating 6DoF pose of a solar panel, according to some embodiments. [Figure 59]1 is a schematic diagram of a sensor fusion architecture for determining a 6DoF pose of a solar panel, according to some embodiments. [Figure 60] FIG. 1 is a schematic diagram of another sensor fusion architecture for determining 6DoF pose of a solar panel, according to some embodiments. [Figure 61] FIG. 1 is a schematic diagram of a sensor fusion architecture according to some embodiments. [Figure 62] FIG. 1 is a schematic diagram of a solar panel installation according to some embodiments. [Figure 63] 1 is a flow diagram of a method for autonomous solar power installation according to some embodiments. DETAILED DESCRIPTION OF THE INVENTION

[0015] The features and advantages of the present invention will become more apparent from the following detailed description when taken in conjunction with the drawings, in which like reference numbers identify corresponding elements throughout. Generally, in the drawings, like reference numbers indicate identical, functionally similar, and / or structurally similar elements.

[0016] Reference will now be made in detail to the embodiments of the present invention, examples of which are illustrated in the accompanying drawings.

[0017] 1 is a perspective view of a solar panel handling system with a solar panel box according to one embodiment of the present disclosure. The solar panel handling system may include an end of arm assembly tool 100 that can couple to individual solar panels 120 from the solar panel box and move the panels to an installation position relative to an installation structure.

[0018] The end of the arm assembly tool 100 may include a frame 102 and one or more mounting devices 104 coupled to the frame 102. Some example mounting devices 104 include suction cups or other structures capable of releasably attaching to the surface of the solar panel 120 and maintaining the attachment, at least collectively, during manipulation of the solar panel 120 by the end of the arm assembly tool 100. The frame 102 may be comprised of multiple trusses 102-A to provide structural strength and stability to the frame 102. Additionally, the frame 102 also serves as a base for the end of the arm assembly tool 100 and other associated components of the solar panel handling system disclosed herein.

[0019] Other relevant components of the solar panel handling system disclosed herein may be coupled to frame 102 to fix the relative positions of these components on the end of arm assembly tool 100. One or more of the various components of the solar panel handling system may be coupled to one or more of trusses 102-A to fix the relative positions of the components on the end of arm assembly tool 100.

[0020] The attachment device 104 is configured to securely attach to a flat surface, such as the surface of a solar panel, for example, by using a vacuum. In one suction cup embodiment, the suction cup can be activated by pressing the suction cup against the flat surface, which forces air out of the suction cup and creates a vacuum seal against the flat surface. As a result, the flat surface adheres to the suction cup with an adhesive strength that depends on the size of the suction cup and the integrity of the seal to the flat surface. In some embodiments, the suction cup engages the solar panel to form an airtight seal, and then a vacuum pump draws air out of the suction cup, thereby creating the vacuum required for proper attachment to the solar panel. In some embodiments, when the flat surface is sealed to the suction cup, an air inlet (not shown) supplies air to the flat surface, thereby releasing the vacuum and releasing the flat surface from the suction cup.

[0021] The system may further include a linear guide assembly 106 coupled to the end of the arm assembly tool 100. The linear guide assembly 106 includes a linearly movable clamping tool 108 having an engagement member 108-A configured to engage a clamping assembly coupled to the installation structure. The linear guide assembly 106 may be actuated to move the clamping tool 108 along an axis, such as between an extended position and a retracted position. The axis of movement of the clamping tool 108 may be parallel to the axis of the installation structure. Thus, the linear guide assembly 106 is capable of moving the clamping tool 108 and the engagement member 108-A along the installation structure.

[0022] In some embodiments, the engaging member 108-A may comprise an electromagnet that can be actuated to grip the clamping assembly 602 (see FIGS. 6A, 6B). Alternatively or additionally, the engaging member 108-A may comprise a gripper to prevent disengagement between the clamping assembly 602 and the engaging member 108-A when the linear guide assembly 106 is actuated to move the clamping tool relative to the installation structure, as described in more detail elsewhere herein.

[0023] The linear guide assembly 106 is actuated using a force-torque transducer 110. In some embodiments, the linear guide assembly 106 and the force-torque transducer 110 may form a rack-and-pinion arrangement, in which case rotation of the force-torque transducer 110 results in the clamping tool 108 moving forward or backward. In some embodiments, the linear guide assembly 106 may be a hydraulic assembly including a telescoping shaft coupled to the clamping tool 108. In such embodiments, the force-torque transducer 110 may be in the form of a pump for pumping hydraulic fluid. In other embodiments, the force-torque transducer 110 may be in the form of or coupled to a linear drive motor that engages a surface of a telescoping shaft coupled to the clamping tool 108.

[0024] In some embodiments, the linear guide assembly 106 may include an electric rod actuator for moving the clamping tool 108 parallel to the axis of the installation structure.

[0025] In some embodiments, the guide assembly 106 may include rollers 606 to facilitate movement of the clamping tool 108 along the mounting structure 604. The rollers may include, for example, bearings or other components designed to reduce friction as the clamping tool 108 moves relative to the mounting structure. The rollers may be coupled to sensors, such as force or rotation sensors, to provide feedback to the controller.

[0026] In some embodiments, the guide assembly may include a spring mechanism 608 that allows for a small amount of tilt (up to 15 degrees) of the clamp tool 108 relative to the installation structure 604. Such tilt may occur when the orienting assembly 804 tilts the end of the arm assembly tool 100 relative to the installation structure 604 to properly level the solar panel.

[0027] The system may further include a junction box 112 coupled to the frame 102. The junction box 112 may include a controller configured to control the force-torque transducer 110 and the mounting device 104. In some embodiments, the junction box 112 may further include a power supply or power controller for controlling power to the various components.

[0028] In some embodiments, the controller 112 may include a processor operatively coupled to memory. The controller 112 may receive inputs from sensors associated with the solar panel handling system (such as, for example, a light sensor or proximity sensor 108-B described elsewhere herein). The controller 112 may then process the received signals and output control commands to control one or more components (such as, for example, the linear guide assembly 106, the clamping tool 108, or the mounting device 104). For example, in some embodiments, the controller 112 may receive a signal from a proximity sensor that determines that the clamping assembly is approaching the trailing edge of the solar panel being installed, and in response, reduce the speed of the linear guide assembly 106 to reduce excessive force and impact on the solar panel.

[0029] Referring to FIG. 8 , in some embodiments, the solar panel handling system may further include an optical sensor 802, such as a camera, a photodetector, or any other optical imaging or light-sensing device. The optical sensor may be suitably positioned on the frame 102, such as on the outer or lower surface of an edge member, as shown by position 802-A in FIG. 8 , or at an interior position of the frame 102 having a field of view that includes the front edge of the solar panel, as shown by position 802-B in FIG. 8 . The optical sensor may be configured to sense the orientation of the solar panel relative to the installation structure during operation of the arm assembly tool end. In some embodiments, the optical sensor may be configured in the form of one or more optical guidance levels (not shown). In such embodiments, one or more light beams (e.g., laser beams) may be projected from one end of the arm assembly tool 100 end, such as a first position on the frame 102, along or parallel to the axis of the installation structure 604. One or more photodetectors may be positioned at another end of the arm assembly tool 100 end, such as a second position on the frame 102, to detect the one or more laser beams. Thus, if the solar panel 120 being installed is not properly oriented or properly leveled relative to the installation structure 604, the solar panel 120 may block some or all of the one or more laser beams, resulting in a change in the signal from one or more photodetectors, which may indicate that the solar panel 120 is not properly oriented or properly leveled relative to the installation structure 604.

[0030] In some embodiments, one or more sensors, such as optical sensor 802, may be used to detect and recognize objects for improved accuracy in positioning and control of installation. These sensors may be implemented with a neural network, such as an artificial intelligence (AI) system. For example, the neural network may include acquiring and modifying images related to the solar panel handling system, solar panels (both existing and to-be-installed), and the installation environment (both the natural environment, such as terrain, and existing equipment, such as structures associated with the solar panel array). Further, for example, the neural network may include acquiring and modifying position or proximity information. The modified images and / or modified position or proximity information are input into the neural network and processed to estimate movement and position of equipment in the solar panel handling system, such as those associated with autonomous vehicles, storage vehicles, robotic equipment, and installation equipment. The estimated movement and position are published to control systems associated with individual pieces of equipment in the solar panel handling system or to a master controller for the entire solar panel handling system.

[0031] In some embodiments, a signal from the optical sensor may be input to a controller. In some embodiments, the solar panel handling system may further include an orienting assembly 804 (see FIG. 8 ) configured to tilt the end of the arm assembly tool 100 relative to the installation structure 604. In such an embodiment, the controller 112 may control the orientation in response to input from the optical signal that indicates that the solar panel being installed is not properly oriented or properly leveled relative to the installation structure, such as the torque tube 604. Although the orienting assembly 804 is illustrated as being coupled to the force-torque transducer 110, other means of implementing the orienting assembly 804 will be readily apparent to those skilled in the art.

[0032] In some embodiments, the controller 112 may be configured to control the attachment device 104 to activate or deactivate the attachment / detachment of the attachment device 104. In embodiments where the attachment device 104 is a suction cup, a vacuum may enable coupling or detachment of the solar panel 120 to the end of the arm assembly tool 100.

[0033] In some embodiments, the mounting structure 604 may have an octagonal cross-section, such as shown in Figures 6A, 6B, and 7A-7D, which may form a torque tube that prevents unintentional sliding of the clamp assembly 602. However, other cross-sectional shapes may be used, such as, for example, a square, oval, or other shape. Additionally, the mounting structure 604 may use a circular cross-sectional shape.

[0034] In some embodiments, the assembly tool 100 may be configured to couple to an assembly mobile robot 903 (an example of which is shown in FIGS. 9 and 10 ). The assembly mobile robot 903 may be configured to position the end of the arm assembly tool 100 relative to a stack or storage container 905 of solar panels, move a selected solar panel, and position the selected solar panel relative to the installation structure 604. In some embodiments, the assembly mobile robot 903 may be operatively coupled to the end of the arm assembly tool 100 via a force transducer 110 (or an orienting assembly 804, if applicable). In some embodiments, the assembly mobile robot may further be operatively coupled to a controller, allowing an operator of the assembly mobile robot to control various functions of the end of the arm assembly tool 100, such as, for example, actuating and / or deactuating the attachment device 104, advancing and / or retracting the clamping tool, and / or actuating and / or deactuating an engagement member relative to the clamping assembly.

[0035] 1 , 6A, 6B, 7A-7D, 9, and 10, in operation, a solar panel 120 is grasped and positioned on the mounting structure 604. The solar panel is then tilted relative to the mounting structure 604 so that the leading edge of the solar panel (i.e., the edge that will be adjacent to the edge of a previously installed solar panel, or in the case of a first solar panel, the edge that will be adjacent a stop secured to the mounting structure 604) is oriented closer to the mounting structure 604 than the opposite trailing edge. The leading edge is then placed into a receiving channel (either a receiving channel positioned along the edge of a previously installed solar panel, either as part of a clamp assembly or in a stop) and the tilt of the solar panel is reduced to an installed position on the mounting structure. The tilt angle is reduced so that, while the solar panel is urged into the receiving channel, the edge region of the top flat surface of the solar panel (i.e., the photovoltaic active surface oriented toward the sun) is captured within the receiving channel in the installed position. An example of a receiving channel 610 on the clamp assembly 602 is shown in Figures 6A and 6B.

[0036] Once the solar panel is in position on the installation structure, the force-torque actuator 110 actuates the guide assembly 106 on the end of the arm assembly tool 100 to contact the engaging member 108-A of the clamping tool 108 with the clamping assembly 602. This clamping assembly was originally positioned outside the area on the installation structure that the solar panel would occupy, but close enough for the associated components on the end of the arm assembly tool 100 to reach. The surfaces and features of the engaging member 108-A may be positioned and sized to match complementary features on the clamping assembly 602. After this contact, the force-torque actuator 110 is actuated (either by continuing to actuate or by actuating in a second mode) to slide the clamping assembly 602 axially over a portion of the length of the installation structure 604. The axial sliding of the clamping assembly 602 engages the receiving channel of the clamping assembly 602 with the rear end of the most recently installed solar panel. A sensor located within the force-torque actuator 110 or clamping tool 108, etc., can provide feedback to the controller indicating when the receiving channel of the clamping assembly 602 has fully engaged the rear end of the solar panel. Once the clamping assembly 602 is positioned, the guide assembly 106 retracts, allowing the next solar panel installation to occur.

[0037] In some embodiments, the linear guide assembly 106 may include a proximity sensor 108-B configured to sense the distance between the engagement member 108 and the rear edge of the solar panel 120 during installation of the solar panel 120. The output from this proximity sensor 108-B may be used to avoid excessive force and impact on the solar panel 120 by appropriately controlling the speed of the clamping tool 108 during operation of the linear guide assembly 106. In some embodiments, the proximity sensor 108-B may be, for example, an optical sensor or an audio sensor (e.g., sonar) that detects the distance between the front edge of the solar panel 120 and the engagement member 108. In other embodiments, the proximity sensor 108-B may be a limit switch that is retracted upon contact.

[0038] 9 and 10 , the assembly mobile robot 903 may be implemented using a land vehicle 907. For example, the land vehicle 907 may be implemented as an electric vehicle (EV). The land vehicle 907 may move autonomously in the vicinity of the installation structure 604. Although not shown, the land vehicle 907 may move along a track or rail attached to or spaced apart from the installation structure. In some embodiments, the land vehicle 907 may be controlled using sensors or based on input or feedback from sensors. These sensors may be, for example, optical sensors or proximity sensors. In further embodiments, a neural network using artificial intelligence may be used to control the movement of the land vehicle 907, for example, by analyzing the operating environment and developing instructions for the movement of the land vehicle.

[0039] FIG. 10 shows an embodiment of a solar panel handling system with two robotic arms, where two assembly tools are coupled to an assembly mobile robot using their respective robotic arms.

[0040] As shown in FIG. 9 , a storage container 905 containing the solar panels to be installed may be disposed on a land vehicle. Here, FIG. 9 illustrates a solar panel handling system including an arm assembly tool 100 coupled to an assembly mobile robot using a robotic arm. Alternatively, as shown in FIG. 10 , one or more storage containers 905 may be disposed on each one or more of the modular vehicles 1005 adjacent to the land vehicle 907. Thus, FIG. 10 illustrates a solar panel handling system with two robotic arms, in which case two assembly tools are coupled to the assembly mobile robot using respective robotic arms. In embodiments of the present disclosure, these robotic arms may be articulated arms having two or more sections connected by joints, or alternatively, may be truss arms. The drawings herein are intended to disclose the use of any type of arm according to the present disclosure.

[0041] Referring to FIG. 9 , a robotic arm, such as arm assembly tool 100 having upper and lower sections 908 and 909, can increase freedom of movement while maintaining light weight and operational simplicity. As further shown in FIG. 9 , a second robotic arm 911 may include arm assembly tool 100 having a nut runner or nut driver at its end for fastening the solar panel to installation structure 604. While any type of robotic arm may be used for second robotic arm 911, FIG. 9 illustrates the use of an articulated arm having a nut runner or nut driver at its end. In this case, robotic arms 100 and 911 may operate autonomously using computer vision with neural networks and artificial intelligence control. Alternatively, robotic arms 100 and 911 may be manually or remotely operated.

[0042] In some embodiments, the land vehicle 907 may be an autonomous vehicle, where neural networks and artificial intelligence control movement and operation, and the modular vehicles 1005 are towed or coupled to the land vehicle 907. In other embodiments, the modular vehicles 1005 may be autonomous vehicles, where neural networks and artificial intelligence control movement and operation, and the land vehicle 907 is towed or coupled to the modular vehicles 1005. Additionally, in some embodiments, the assembly mobile robot 903 is mounted on one of the land vehicle 907 and the modular vehicles 1005. In other embodiments, the assembly mobile robot 903 may be mounted on a dedicated robotic vehicle.

[0043] A process for installing solar panels is shown in FIGS. 11A-11C. As shown in FIG. 11A, pallets of solar panels may be delivered by truck. In some embodiments, the pallets may constitute storage containers 905 for the solar panels. The pallets may include machine-readable symbols, such as barcodes, QR codes, or other manufacturing references, that can be read to provide information about the solar panels, installation procedures, or other information used in the installation process, particularly information used by neural networks and artificial intelligence control. Such information may include, for example, the number of solar panels, the type of solar panel, physical characteristics such as the size of the solar panels, installation-related characteristics such as the type and location of hardware, installation procedures, or other characteristics of the solar panels, storage of the solar panels on the pallets, and information related to installation. Furthermore, by using the machine-readable symbols, the system can control the supply or replenishment of panel boxes in the correct order and / or ensure that panels with similar impedance from the factory are used.

[0044] As shown in FIG. 11B, a mechanized device such as a forklift may be used to move and position the pallet on the land vehicle. In this case, the forklift may be manually operated, remotely operated, or autonomous. In FIG. 11B, the pallet is positioned on the land vehicle. Alternatively, the pallet may be positioned on a modular vehicle. Then, as shown in FIG. 11C, a robotic arm is used to install the solar panels. In the illustrated example, two arms are used to handle each solar panel to be installed on each installation structure. In this case, the land vehicle moves between the two installation structures. Additionally, one modular vehicle is provided, but this may be separate from the land vehicle.

[0045] As will be appreciated by those skilled in the art, modifications and variations of the implementation may be made. For example, as shown in Figures 12A and 12B, two modular vehicles may be provided for each robotic arm. In a further alternative, the modular vehicles may be coupled to the land vehicle without being spaced apart. Thus, as shown in Figure 12A, the robotic arm may engage with each solar panel to be installed, as shown in Figure 12B.

[0046] In some embodiments, placement may be achieved using computer vision registration, as shown in Figure 13. For example, as described above, optical sensors and the like may be used in conjunction with neural networks for artificial intelligence.

[0047] 14, in some embodiments, when modular vehicles are used with land vehicles, the modular vehicles may be replaced with replenished modular vehicles when all of the modular vehicle's solar panels have been installed. In this case, computer vision processes may be used to communicate with and control an autonomous stand-alone vehicle, such as a forklift, to deliver additional solar panel boxes. In this manner, the supply of solar panels may be replenished.

[0048] In a replenishment operation using the forklift example, a forklift (whether autonomous, remotely controlled, or manually operated) can be used to return empty boxes or containers of solar panels to a disposal area, remove straps from boxes being delivered, open lids or cut off box faces, pick up boxes to correct the rotation / orientation of the solar panels, or perform other tasks. Additionally, the forklift may remain near the land vehicle to wait for the system to empty the next box of solar panels. Thus, the forklift can manually or autonomously discard the emptied box, position the next box on the land vehicle or modular vehicle, unpack the box (including removing straps, opening lids, or cutting off box faces), and back away from the land vehicle / modular vehicle. As noted above, this replenishment may be, for example, autonomous, remotely controlled, or manually operated.

[0049] 15 to 34 show detailed diagrams of an example of the configuration of a system for installing solar panels according to an embodiment of the present disclosure.

[0050] Computer Vision and Artificial Intelligence (AI) Technology According to some embodiments, computer vision and AI techniques can be used to determine possible locations for installing solar panels along a mounting structure, such as a torque tube. The installation location of a panel can be based on the location of a previously installed panel. In some embodiments, the six-degrees-of-freedom (6DoF) pose of the previously installed panel is used. The term "6DoF" refers to six degrees of freedom: three rotational axes (yaw, pitch, and roll) and three translational axes (x, y, and z). The estimation of the 6DoF pose of the panel includes the location in 3D space (x, y, z) of a point (or keypoint) on the panel, such as a corner of a frame or some other visual reference, and the rotation angle of the panel around three orthogonal axes (rx, ry, rz). In some embodiments, computer vision or AI is used to estimate the 6DoF pose of the previously installed panel to determine the location where the next panel should be installed along the torque tube. Typically, the next panel to be installed along the torque tube is offset by a fixed amount from the keypoint position (x, y, z) of the previously installed panel and has the same rotation angle (rx, ry, rz).

[0051] 6DoF Pose Estimation for Panel Placement and Panel Picking Although some embodiments described herein relate to panel placement (e.g., for installation along a torque tube), the embodiments rely on techniques that are also applicable to panel picking from a storage location, such as a module box or cradle. In the case of panel picking, the 6 DoF pose of the panel is determined, as in the case of panel placement. In the case of panel placement, the panel may be a previously installed panel, as discussed above. In contrast, in the case of panel picking, the panel may be the current (or outermost) panel (in a storage location) to be picked for installation. In the case of panel placement, the 6 DoF pose of the previously installed panel may be used to determine the (offset) position of the next panel to be installed along the torque tube, as discussed above. In contrast, in the case of panel picking, the 6 DoF pose of the current panel to be installed may be used to determine the pick point (usually the center) of the panel to be picked. In the case of panel picking, the challenge is to pick the panel so that the EOAT is centered relative to the panel. This centering of the EOAT relative to the picked panel ensures an even load distribution on the EOAT. In the case of panel placement, as discussed above, the challenge is to place the next panel with a specific pose (typically x, y, panel width, and [e], z, rx, ry, rz) relative to the previously placed panel. However, for both panel placement and panel picking, the goal of CV / AI is to determine the 6DoF pose of the panel. Furthermore, this 6DoF pose can be determined by methods and techniques based on the methods described herein.

[0052] In the following description, the term "computer vision" will be used to refer to "classical" computer vision techniques and algorithms, and the term "artificial intelligence (AI)" will be used as an alternative to "machine learning (ML)" or "deep learning (DL)" so that this disclosure can be better generalized to current and newly developed technologies.

[0053] The term "image" is used to include not only camera images but also data and images from other types of sensors, such as time-of-flight sensors, LiDAR, or other sensors that image or scan a field of view (FoV). Also, the terms "pre-processing" and "post-processing" may involve different hardware (e.g., computing devices, sensors, etc.) for initial or subsequent processing, respectively.

[0054] Examples of non-AI methods for pose estimation In some embodiments, the 6DoF pose is estimated using non-AI techniques. The 6DoF pose of the panel may be determined using various sensors (e.g., stereo cameras, LiDAR, etc.) and computer vision algorithms. These various approaches are described below as various pre-processing and post-processing methods for AI processing. These pre-processing and post-processing steps may be combined to form a computer vision pipeline that can be used on its own (without AI) to determine the 6DoF pose of the panel.

[0055] AI Model and Inference Examples We then describe different types of AI models and inference that can be applied to determining the 6DoF pose of solar panels. AI-based inference can be based on bounding boxes, segmentation, keypoints, depth, and / or 6DoF pose. We then describe each of these different approaches in turn.

[0056] segmentation Segmentation in AI refers to the process of dividing an image or video into meaningful and semantically coherent regions. The goal of segmentation is to partition visual data into distinct regions based on common characteristics, such as color, texture, or object boundaries. In computer vision, traditional segmentation techniques include methods such as thresholding, region growing, edge-based segmentation, and clustering algorithms, such as k-means or mean shift. These methods rely on manual features and heuristics to segment images.

[0057] In AI, deep learning methods, especially convolutional neural networks (CNNs), have shown outstanding performance in segmentation tasks. Fully convolutional networks (FCNs), U-Net, Mask R-CNN, and DeepLab are popular architectures for semantic and instance segmentation. These models leverage their ability to learn and extract complex features from images, enabling accurate and efficient segmentation.

[0058] Various types of segmentation techniques are available in AI, including semantic segmentation and instance segmentation. Semantic segmentation involves labeling each pixel in an image or video frame with a corresponding class label. The output is a pixel-by-pixel classification map, where each pixel is assigned a semantic category or class label. Semantic segmentation focuses on capturing the semantic meaning of a scene and is used for scene analysis, object recognition, and high-level understanding. Instance segmentation goes beyond semantic segmentation to isolate and identify individual objects in an image. In instance segmentation, a unique label or identifier is assigned to each pixel belonging to a specific object instance. In instance segmentation, each object is individually segmented, allowing for accurate delineation and isolation of object boundaries. This technique is used for object detection, tracking, counting, and detailed object analysis. In the following discussion, the term "segmentation" refers to instance segmentation (unless otherwise specified).

[0059] In some embodiments, the segmentation is utilized to determine the 6DoF pose of the panel. Using the panel segmentation as input, various computer vision algorithms can find the pixel locations in the image of the four corners of the panel frame, and these locations are used as input to a Perspective-n-Point (PnP) solver to determine the 6DoF of the panel in 3D space (or "world space").

[0060] In some embodiments, pixel locations of the panel's corners are determined based on the segmentation by one or more of a variety of computer vision techniques, such as the Hough transform, the Ramer-Douglas-Peucker algorithm, or some other computer vision post-processing. For example, the Hough transform can generate four Hough lines that fit the four edges of the segmentation, and these Hough lines correspond to the four edges of the panel. The four intersections of the four Hough lines generate the four corners of the panel. FIG. 51 illustrates an example application 5100 of the Hough transform to identify corners of a solar panel 5110, according to some embodiments. The intersection of Hough lines 5102 and 5106 generates corner A, the intersection of Hough lines 5102 and 5108 generates corner D, the intersection of Hough lines 5104 and 5106 generates corner B, and the intersection of Hough lines 5104 and 5108 generates corner C.

[0061] In some embodiments, the Ramer-Douglas-Peucker algorithm is used to determine the pixel locations of the four panel corners. In some embodiments, this algorithm is constrained to generate a four-sided polygonal approximation. For example, in an OpenCV implementation of this algorithm, i.e., the approxPolyDP() function, it is possible to generate a four-sided polygonal approximation by optimizing the parameter epsilon, for example, based on a binary search of epsilon. Furthermore, other applicable computer vision algorithms, as recognized by those skilled in the art, may be used for post-processing of the segmentation to find panel corners or other keypoints. In some embodiments, keypoints such as panel corners are used as inputs to a perspective n-point solver to determine the 6DoF pose of the panel.

[0062] In some examples, the above-mentioned post-processing methods (e.g., Hough transform, Ramer-Douglas-Peucker algorithm, etc.) that fit a single line along the longitudinal direction of the panel edge may produce inaccurate corner locations in either the far half or near half of the panel. This is because solar panels are not perfectly rigid. For example, when mounted on a torque tube or the like, the panel is cantilevered on the tube and flexes or deforms under its own weight. Therefore, approximating the longitudinal edge with a single line may be inaccurate for either corner. To compensate for this inaccuracy, in some embodiments, instead of fitting a single line, two lines are fit to the longitudinal edge of the panel, one for the far half and the other for the near half. A perspective n-point solver can then solve the 6DoF pose based on the coordinates (x, y, z) of the panel corners that deviate from perfectly rigid body coordinates (e.g., with z as the height change due to panel deformation).

[0063] Another post-processing method that may be particularly well suited for AI-based segmentation is analyzing pixel intensities in shadows between adjacent panels. AI-based segmentation may perform worse for images containing multiple panels (compared to images containing a single panel). For example, in images containing multiple panels, the segmentation of the panel may recede into the interior of the panel (rather than aligning with the outer edge of the panel frame, as would be expected).

[0064] Such discrepancies may be observed along the length of the panel near its adjacent panels. In some embodiments, to compensate for this discrepancy, the shadow between the panel and its adjacent panels serves as a visual reference for identifying the true edge of the panel. The shadow between adjacent panels is generally consistent and distinguishable across various lighting conditions. In some embodiments, the true edge of the panel is identified using computer vision algorithms, such as binarization, thresholding, or a Hough transform.

[0065] Depth estimation In some embodiments, depth estimation is utilized to determine the 6DoF pose of the panel. Depth estimation is the determination of the distance of a given object within the field of view (FoV) relative to a sensor (e.g., camera, LiDAR, time-of-flight sensor, etc.).

[0066] Depth estimation can be understood as a type of segmentation. The object of interest is segmented or distinguished from the more distant background. Thus, the panel appears as a continuous segmentation (in a disparity map (using computer vision) or depth visualization (using AI)). Figure 52 shows an example application 5200 of segmentation according to some embodiments. The result of this segmentation can be used as input for post-processing to find panel corners or other keypoints for the purpose of 6DoF pose estimation, as described above.

[0067] When using AI, depth estimation can be performed using, for example, neural networks for monocular depth estimation (MDE). For a comprehensive survey of prior art approaches to MDE, see the following reference, which is incorporated herein by reference: J. Spencer et al., "The monocular depth estimation challenge," Proceedings of the IEEE / CVF Winter Conference on Applications of Computer Vision, pp. 623-632, (2023).

[0068] In computer vision applications, depth estimation can be performed using two cameras (or stereo cameras) to generate a disparity map based on corresponding points and the baseline distance between the two cameras. The disparity or perspective difference between images from the two (stereo) cameras indicates the distance of an object from the camera. When using CV, depth estimation becomes uncertain under certain lighting or surface conditions, such as textureless areas and specular reflections on panels. Conventional techniques, such as photoconsistency methods and active stereo (using IR structured light projection), can improve the accuracy of depth estimation. Furthermore, multi-baseline stereo, involving three or more cameras with, for example, a trinocular or tetranocular configuration (where multiple baseline distances exist between the cameras), and related algorithms (such as semi-global matching and iterative matching algorithms), can improve the resolution and accuracy of depth estimation. For a comprehensive survey of prior art approaches to multi-baseline stereo, see the following documents, which are incorporated herein by reference: H. Hirschmuller, “Stereo Processing by Semiglobal Matching and Mutual Information,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 30, no. tMGM) for Dense Stereo Matching”, ArXiv, abs / 1911.09800 (2019).

[0069] Whether using AI or CV, depth estimation (a type of segmentation) utilizes the same post-processing (as described above) to locate panel corners in the image. For example, the determination of corner locations, whether from depth estimation or segmentation, can be modified or refined (before being used as input to the PnP solver) by what is called "two-pass homography" post-processing, described below.

[0070] Keypoint detection Some embodiments determine the 6DoF pose of a panel by utilizing a type of AI-based inference called keypoint detection. Keypoint detection is a fundamental task in computer vision and AI, involving identifying and locating specific points or landmarks in an image or video. These keypoints correspond to distinctive features in the visual data, such as corners, edges, and other regions of interest. The goal of keypoint detection is to accurately locate and describe these keypoints, enabling a variety of applications, such as object recognition, tracking, pose estimation, and image registration.

[0071] Keypoint detection typically involves the following steps or algorithms: Preprocessing: Typically, input images or video frames are preprocessed to improve quality and reduce noise. Common preprocessing steps include resizing, normalization, and grayscale conversion. Depending on the specific requirements and characteristics of the data, various algorithms can be used for keypoint detection. Some embodiments utilize corner detection, i.e., identifying corners in an image based on local intensity variations by utilizing algorithms such as the Harris corner detector or the Shi-Tomasi corner detector. Some embodiments utilize scale-space extremum detection, including methods such as Difference of Gaussians (DoG) or Laplacian of Gaussians (LoG), to detect keypoints at various scales by looking for local extrema in the scale-space representation of the image. Some embodiments utilize interest point detectors, including algorithms such as SIFT (Scale Invariant Feature Transform) or SURF (Superfast Robust Features), to identify keypoints based on local image gradients and their response to various scales and orientations. Some embodiments utilize deep learning-based methods such as convolutional neural networks (CNNs), which can be trained to detect keypoints directly, for example, by learning from annotated datasets. Models such as DenseNet, CornerNet, and OpenPose utilize deep learning techniques for keypoint detection. Once a keypoint is detected, its exact location within the image must be determined. This process may involve using techniques such as subpixel interpolation or optimization algorithms to refine the initial detection and improve localization accuracy.

[0072] The key points may be the four corners of the solar panel frame as discussed above, or any visual reference provided by a single panel, such as a common grid pattern, as discussed below, or a point on the solar panel, such as its center.

[0073] As with segmentation models / inference, keypoint detection models / inference can utilize unique computer vision post-processing methods. For example, a COSFIRE (Combination Of Shifted Filter Responses) filter can be optimized for corner pattern recognition to distinguish corners from the background. The keypoint detection model generates a coarse region of interest (ROI) to input into the COSFIRE filter, which can result in finer corner detection.

[0074] Other applicable computer vision algorithms and techniques as will be appreciated by those skilled in the art may be utilized to post-process the keypoint detection to find panel corners for the purpose of 6DoF pose estimation.

[0075] Bounding Box In some embodiments, the 6DoF pose of the panel is determined by using bounding boxes, another AI-based inference technique. Bounding box inference in AI refers to the process of predicting and locating objects in an image or video by drawing a rectangular bounding box around them. Bounding boxes provide an approximation of the object's location and extent within visual data. This technique is widely used in object detection, localization, and tracking tasks. Bounding box inference typically relies on an object detection model trained using a machine learning algorithm. Common object detection models include the Single Shot Multibox Detector (SSD), You Only Look Once (YOLO), and the Faster R-CNN (Region-based Convolutional Neural Network). These models are designed to detect and classify objects within images or video frames. Object detection models are trained on a labeled dataset, where each object of interest is annotated with a bounding box that tightly surrounds the object. Additionally, the training data includes class labels corresponding to each object category. During training, the model learns to identify objects and predict bounding boxes for those objects based on visual features extracted from the input data. During the inference phase, the trained object detection model receives input images or video frames as input and processes the input to detect objects and predict their bounding boxes. The model analyzes the visual features and makes predictions about the presence, class, and location of objects in the image. The object detection model predicts bounding box coordinates for each detected object.Typically, a bounding box is represented by four numbers: the x and y coordinates of the upper-left corner, and the width and height of the box. These numbers are used to draw a rectangle around the object to suggest its estimated location within the image or frame. Because multiple bounding box predictions may overlap or enclose the same object, a post-processing step called non-maximum suppression is often applied. Non-maximum suppression aims to ensure that only the most accurate and reliable bounding box remains for each object by removing redundant or overlapping bounding boxes. This step helps eliminate duplicate detections and improve the accuracy of bounding box inference.

[0076] In some embodiments, a bounding box is used to identify a given solar panel by bounding, or enclosing, the entire solar panel in an image. A bounding box can be understood as a special case of keypoint detection aimed at determining the panel's 6DoF pose. This is because, given a typical top-down perspective view, two corners of the bounding box may coincide with two physical corners of the panel frame. Specifically, the upper-right and lower-left corners of the bounding box coincide with the rightmost and leftmost corners of the panel frame. This correspondence allows bounding box inference to be interpreted as keypoint detection (of two corners). This correspondence allows the keypoint detection model to be replaced with a bounding box model in approaches that do not require identification of all four panel corners, which may be advantageous for training and / or inference.

[0077] Whether as keypoint detection or bounding boxes, AI-based inference for determining keypoints, such as the corners of a panel, can utilize the same post-processing (such as the "two-pass homography" algorithm described below) for the purpose of determining the 6DoF pose of the panel.

[0078] 6 degrees of freedom Six-degrees-of-freedom (6DoF) pose estimation refers to the task of estimating the position and orientation of an object in 3D space using machine learning models. The term "6DoF" stands for six degrees of freedom, which correspond to three rotational axes (yaw, pitch, and roll) and three translational axes (x, y, and z). Pose estimation is essential in a variety of computer vision applications, such as robotics, augmented reality, and object tracking, where accurately knowing the position and orientation of an object is crucial.

[0079] Some embodiments utilize AI models for 6DoF pose estimation. These AI models can be categorized into two main types: feature-based methods and direct regression methods.

[0080] First, feature-based methods extract keypoints or features from an object or scene and match them between a 3D model and an input image. Then, a pose is estimated based on the spatial relationship between the matched 3D-2D feature correspondences. Feature extraction and matching can be performed using traditional computer vision techniques or deep learning-based methods. Traditional feature-based methods use techniques such as SIFT (Scale Invariant Feature Transform), ORB (Orientation Fast and Rotation Brief), or SURF (Speed ​​Robust Features) to detect and match keypoints between a 3D model and a 2D image. Then, RANSAC (Random Sample Consensus) or PnP (Perspective n-Point) algorithms are used to estimate a 6DoF pose from the matched correspondences. Deep learning-based feature-based methods can learn feature representations directly from input images using deep learning models such as CNNs instead of using manually designed features. PoseNet and PoseCNN are examples of deep learning-based feature-based methods that use CNNs to predict 6DoF poses.

[0081] Second, direct regression methods directly predict 6DoF pose parameters (translation and rotation) from input images, avoiding the need for feature extraction and matching. CNN-based regression models directly regress 6DoF pose parameters from input images by using a CNN architecture. The network receives an image as input and outputs pose values ​​as a series of numbers. Models such as DeepIM and PVNet are examples of CNN-based regression methods. Some hybrid approaches combine feature-based and direct regression methods. These hybrid approaches achieve higher accuracy by utilizing a CNN to predict an initial pose estimate and then utilizing feature-based methods to refine the pose.

[0082] If AI models for 6DoF inference have sufficient accuracy, they can be used to determine the 6DoF pose of solar panels (without post-processing corrections). However, even if the accuracy of 6DoF models is not sufficient, these models can still provide information that can be interpreted in terms of other types of inference, such as segmentation, keypoint detection, and bounding boxes, which can then be corrected or refined through post-processing. For example, an initial (rough) 6DoF estimate of panel pose can be used to roughly identify corners for the purposes of keypoint detection or bounding box inference. The initial 6DoF pose estimate can also serve, for example, to (roughly) identify or segment the panel from its background.

[0083] In summary, what has been described above are various types of AI models and inferences that can be applied to determine the 6DoF pose of solar panels. The AI ​​models can be zero-shot, one-shot, or few-shot, and can be trained in different ways, for example, using real or synthetic images, or using supervised or unsupervised learning.

[0084] In some embodiments, each AI model or inference utilizes a unique post-processing method (e.g., analysis of shadows between adjacent panels, filters, etc.) In some embodiments, different AI models / inferences utilize the same post-processing method (e.g., "two-pass homography," described below).

[0085] These AI models may be combined, with the output of one model being used as input to another. For example, a first model (such as YOLO) outputs bounding boxes, which are then used as input to a second model (such as YOLO, SSD, RCNN, or SAM) for segmentation. A model of one type may be used twice, with the second instance of the same model receiving the output of the first as input. For example, YOLO for bounding boxes may be used as input to YOLO for keypoint detection. This AI pipeline may consist of any number of AI models or ensembles of AI models. In machine learning, ensemble learning refers to the technique of combining multiple individual models, called base models or weak learners, to form a more powerful and accurate model known as an ensemble model. The idea behind ensemble learning is that by aggregating the predictions of multiple models, the ensemble can make more reliable predictions than each model alone. There are several ensemble methods used to combine the predictions of base models. Base models are the individual models that form the ensemble. The base models can be any machine learning algorithm, such as a decision tree, random forest, support vector machine, neural network, or any other model. Each base model is trained on a subset of the training data or with some variation to provide diversity among the models. For example, in a voting-based ensemble, each base model makes a prediction independently, and the final prediction is determined by majority voting (for classification problems) or averaging (for regression problems) among these predictions. Bagging (bootstrap aggregation) involves training multiple base models on different random subsets of the training data with replacement. The final prediction is obtained by averaging (regression) or voting (classification) the predictions of the individual models.Boosting algorithms, such as AdaBoost, Gradient Boosting, or XGBoost, train base models sequentially, with each subsequent model focusing on instances that the previous model struggles with. The predictions of all models are combined to form the final prediction, often through weighted voting. Stacking combines the predictions of multiple base models by training a meta-model that learns to make predictions based on the outputs of the individual models. The predictions of these base models are used as features, and the meta-model is trained on this augmented dataset.

[0086] The strength of an ensemble lies in the diversity among its base models. Diversity can be achieved by various means, such as using different algorithms, changing the model architecture, training on different subsets of data, or introducing randomness during training. Different models generate different types of errors, which, when combined, can compensate for each other's weaknesses, resulting in improved overall performance. Ensembles often outperform individual models because they can capture different aspects of the data and combine their strengths to produce more accurate predictions. Typically, ensembles are more robust to noise and overfitting than individual models, since errors from some models can be compensated for by others. Ensembles can generalize to unknown data by mitigating the effects of bias and errors from individual models, resulting in better performance.

[0087] Described above are similarities or overlaps in application between various types of AI models / inferences, which allow one type of model / inference to be interpreted as a special case of another (more general) type of model / inference. Such generalizability is based on the specific characteristics of the imaged scene in the use case of the present invention. Conversely, these various types of models / inferences can be subjected to the same post-processing (e.g., "two-pass homography" algorithm) based on their generalizability.

[0088] Described below are ways in which various AI models / inferences can be combined with various computer vision algorithms. Computer vision algorithms can be integrated into an AI pipeline as post-processing or pre-processing. Furthermore, computer vision algorithms can be integrated into an AI pipeline based on an architecture that leverages the strengths of computer vision and AI by enabling deeper complementarity between them. Computer vision algorithms can also be combined on their own without using AI.

[0089] Additionally, the deeper complementarity between CV and AI in this disclosure exemplifies what O'Mahony et al. (2020) characterized as "a blend of hand-designed approaches and [deep learning] DL for improved performance." They state, "There is a clear trade-off between traditional CV and deep learning-based approaches. Classical CV algorithms are well established, transparent, and optimized for performance and power efficiency, while DL offers higher accuracy and generality at the expense of significant computing resources." However, they also state that for applications where DL is not yet well established (e.g., 3D vision), it is possible to effectively combine CV and AI / DL. See N. O'Mahony et al., "Deep learning vs. traditional computer vision," Advances in Computer Vision: Proceedings of the 2019 Computer Vision Conference (CVC), Springer Nature Switzerland AG, Volume 11 (pp. 128-144) (2020).

[0090] Whether post-processing or pre-processing, part of deeper integration with AI, or a self-contained computer vision pipeline without AI, computer vision algorithms rely on invariant structures that can be identified across different images. Such invariant structures may belong to physical structures, such as the grid pattern formed by photovoltaic cells and their electrical connections. Furthermore, such invariant structures may belong to non-physical structures of lighting, such as a type of negative space, such as the shadow cast between adjacent solar panels. Whether physical or non-physical, positive or negative space, the invariant structures on which CV algorithms are based provide consistent visual patterns that can be analyzed for semantic content.

[0091] Post-processing using the "two-pass homography" algorithm Post-processing to correct or correct segmentation inaccuracies (AI-based or computer vision-based) is called a "two-pass homography" (TPH) algorithm. This algorithm relies on regular visual patterns or visual references on solar panels, such as grid patterns. The grid pattern (or other visual references) provide reference points that allow the pixel locations of panel corners to be indirectly inferred through a homography transformation, and can therefore be used to correct / correct segmentation inaccuracies. In some embodiments, the TPH algorithm includes the following steps: (1) Filter out the background in the image by using the panel segmentation as a mask (leaving only the foreground panels). (2) Identifying the grid intersections of the masked panel of (1) as "corners" by using the Shi-Tomasi algorithm. (3) Calculating the homography matrix H1 by using the "corners" of (2). (4) Determining the location of the grid intersections of (2) in millimeter space based on H1 of (3) (described below). (5) Transforming the grid intersections of (2) in image space to positions in millimeter space of (4) by determining a homography matrix H2. (6) Back-projecting the corners of the millimeter space into the image space using the inverse matrix of H2 in (5).

[0092] In computer vision parlance, the term "corner" refers (roughly) to a point of interest where the intensity gradient is greatest in all directions. "Corner" (or italic corner) should be distinguished from a physical corner (e.g., the corner of a panel frame).

[0093] Regarding (1), the filtering operation can be as simple as a binary AND between the mask and the raw image.

[0094] Regarding (2), other corner-finding algorithms (e.g., the Harris-Stephens algorithm) may be used. Note that the grid intersections or “corners” found in (2) may not perfectly coincide with the intersections of grid lines. This is because these grid intersections may consist of various shapes (e.g., diamonds, two triangles with a common vertex, etc.) with gradient strengths that result in off-center “corners” as detected by the corner detection algorithm. Figure 53 shows an example (5300). However, the (random) off-centering of the detected “corners” can be counteracted by increasing the number of detected “corners” as input to the homography calculation. In Figure 53, thicker, more noticeable grid lines (sometimes called “major” grid lines) are useful for the TPH algorithm. “Major” grid lines can be distinguished from (thinner) “minor” grid lines by using various computer vision algorithms, such as Gaussian blur with a kernel size of 5.

[0095] Furthermore, in some embodiments, an AI model may be used to find the grid intersections. The AI ​​model may use as input either the raw image or a projected (or warped) image based on the (“first-pass”) homography matrix H1. In H1 space, the panel is projected relative to an approximate top view. This view minimizes perspective variations, making it easier to train the AI ​​model. However, if the AI ​​model receives the raw image as input and generates all grid intersections by inference, then the “two-pass homography” algorithm is replaced by a “one-pass” algorithm. This is because, once all grid intersections are found by the AI ​​model, the correspondence in (4) is replaced by a simple ordering (i.e., the relative positions of the grid intersections are the same), as discussed below.

[0096] Regarding (3), the "first-pass" homography matrix H1 may produce a projection of the panel in "millimeter space" that is not a perfect square. The term "millimeter space" refers to the 2D projection space in which the four (physical) corners of the solar panel have the following (x,y) coordinates: (0,0), (0,L), (W,0), and (W,L), where W and L are the width and length of the solar panel, respectively. However, the homography matrix H1 should be accurate enough to allow for a correspondence (in H1 space) between the grid intersections discovered in (2) and their locations in millimeter space.

[0097] Regarding (4), the correspondence (in H1 space) may be based on Euclidean distance (e.g., a fraction of the shortest cell dimension, such as width) or some other measure. For example, a Euclidean distance that is a fraction of the shortest cell dimension, such as cell width, may ensure that the correspondence between the grid intersections found in (2) and their locations in millimeter space is unique. If the associated Euclidean distance is excessively large (e.g., greater than about half the width or height of an individual cell), a non-unique, ambiguous correspondence or a correspondence suggesting adjacent panels may result.

[0098] For (5), the ("second pass") homography matrix H2 is recalculated based on more accurate (associated) grid-intersection locations in millimeter space as determined in (4).

[0099] For (6), a more accurate homography matrix H2 is used to backproject the corner with coordinates (0,0), (0,L), (W,0), (W,L) in millimeter space into image space, shown by the crosshairs in Figure 53.

[0100] FIG. 54 is a schematic illustration of detecting 5400 grid intersections 5402 using the two-pass homography algorithm described herein, according to some embodiments.

[0101] Some embodiments utilize a form of the TPH algorithm that is generalized to take any panel feature as a visual reference, rather than being limited to grid intersections, as follows: 1. In the image space, the four corners of the panel are k i Estimate. 2.k i Find the homography matrix H1 that transforms from to coordinates in millimeter space, i.e. (0,0), (H,0), (H,W), (0,W). 3.p i and q iwhere p i is the pixel location of the panel feature (as a reference) in image space, and q i is the corresponding position in millimeter space based on H1. 4.p i From q i Find the homography matrix H2 that transforms 5. Inverse matrix H2 of homography matrix H2 -1 Based on this, we backproject the corners (0,0), (H,0), (H,W), (0,W) from millimeter space to image space.

[0102] In the generalized formulation above, p i and q i Note that p can be not only a grid intersection, but also, for example, a diamond centroid or a common vertex of two triangles. i and q i may be some other pixel structure rather than just a "corner".

[0103] Furthermore, in the generalized formulation above, the correspondence between the corresponding source pixel locations and destination pixel locations as shown in (3) and (4) still ultimately relies on the k coordinates of the four corners (not "corners") of the panel as shown in (1) and (2) as the basic reference points. i Note that other standards may be used as reference standards depending on the shape and characteristics of the solar panel.

[0104] Preprocessing using multi-baseline stereo As discussed above, depth estimation can be used for segmentation. Additionally, depth estimation can be used to determine a bounding box that encompasses the entire extent of a panel. The bounding box can then be used to extract or segment a region of interest (ROI) from the raw image for input into an AI pipeline.

[0105] 55 shows a schematic diagram of ROI segmentation 5500 according to some embodiments. This ROI segmentation is performed on a raw (full size) image, such as from two 12MP RGB cameras (5502 and 5504), after which this (segmented) image is sent as input to an AI pipeline 5506.

[0106] As a pre-processing, the ROI segmentation may be based on any of the following: · d1: Disparity map from single baseline stereo. · d2: Disparity maps from multi-baseline stereo. · d3, d4: Monocular depth estimation from neural networks. · f: any (weighted) combination of the above. m: On-board processing that performs bounding box as ROI inference (e.g., processing provided by Luxonis OAK-D and Stereolabs ZED cameras).

[0107] Figure 55 is a diagram of a design space rather than a design. Figure 55 shows the preprocessing design space described above, illustrating a comprehensive design of all elements. In practice, not all of these elements need to be used. This design space considers the structure in motion (including EOAT motion), an ensemble of AI models (h) (Model 1, , Model n), and a distributed Robot Operating System (ROS) architecture (e.g., ROS node i, , ROS node j).

[0108] Sensor fusion using LiDAR and monocular AI Described above are embodiments that utilize computer vision techniques for pre- and post-processing for monocular AI (i.e., using images from a single camera). As will be described below, in some embodiments, computer vision is utilized for more than just a preparatory or correction step in determining the 6DoF pose. Computer vision can be more tightly integrated (or incorporated) into AI.

[0109] Figure 56 illustrates multi-data channel capabilities 5600 of a LiDAR utilized by some embodiments. Distance 5602 is the distance of a point from the LiDAR camera calculated using the time-of-flight of the laser pulse. Signal 5604 is the intensity of the laser return from the object (typically represented by a colored point cloud). Ambient light 5606 is a camera return capturing the intensity of ambient light at a predetermined wavelength (e.g., 865 nm wavelength). Reflectance 5608 is the reflectance of the surface (or object) detected by the LiDAR sensor. Some embodiments utilize a LiDAR that provides access to reflectance, ambient near-infrared (NIR), and (2D projected) range data or images, as well as a LUT mapping between these images. The same pixel sensor may be used for the various data channels. A single LUT or lookup table (capturing the physical dimensions of the hardware) may be sufficient. External calibration may not be required, nor is it necessary to ask whether the calibrated mapping can be generalized. LUT mapping allows analysis to proceed smoothly from one image space to another, for example from NIR images to point clouds.

[0110] Figure 57 is a schematic diagram of an example computer vision / AI architecture 5700 for determining the 6DoF pose of a solar panel, according to some embodiments. The CV / AI pipeline 5700 receives as input two types of data or images from the LiDAR: an ambient NIR grayscale image 5702 and range data in the form of a point cloud 5706. A transform 5704 converts the NIR to range data. First, the ambient NIR image 5702 is converted to the coordinates in image space (x pixel , y pixel The point cloud distance data 5706 is used as input to a CV / AI pipeline (pre-processing 5708 and AI 5710) that localizes the NIR image space coordinates (x, r, rz) of the associated keypoints. This CV / AI pipeline may consist of any of the pre-processing, AI models (e.g., segmentation, keypoint detection, bounding boxes, etc.) and post-processing steps described above. Furthermore, the point cloud distance data 5706 is used as input to a CV pipeline that identifies panels by plane segmentation 5712 (i.e., fitting a plane to the points in the point cloud). The resulting fitting plane, e.g., in a Hessian normal form, yields the rotation angles (rx, ry, rz). Furthermore, the coordinates in NIR image space (x, r, rz) of the associated keypoints are also used. pixel , y pixel ) 5718 is transformed (T) or mapped (5714) into a point cloud to generate keypoint coordinates (x, y, z) in world space. The combined (5716) output (x, y, z, rx, ry, rz) is the 6DoF pose of the panel.

[0111] FIG. 58 is a schematic diagram of another example computer vision / AI pipeline 5800 for determining the 6DoF pose of a solar panel, according to some embodiments. This architecture is similar to that described above with reference to FIG. 57, except that pipeline 5700 relies solely on data or imagery from a LiDAR. In contrast, pipeline 5800 relies on both a LiDAR 5806 and another sensor, i.e., a camera 5802. The operating principles and architecture of pipelines 5700 and 5800 are similar. For example, pre-processing steps 5708 and 5808 are similar, plane segmentation steps 5712 and 5812 are similar, AI steps 5710 and 5810 are similar, and combinations 5716 and 5816 are similar. Similarly, pipeline 5800 achieves combined 6DoF pose estimation by combining point cloud segmentation with AI-based keypoint detection. The only difference is that the CV / AI pipeline shown in Figure 58 utilizes a transformation (T) 5814 rather than a look-up table (LUT) like the pipeline shown in Figure 57. Rather, the transformation or mapping (T) 5814 between the camera and LiDAR image spaces is achieved by calibration.

[0112] Figure 59 is a schematic diagram of a sensor fusion architecture 5900 for determining 6DoF pose of a solar panel according to some embodiments. This architecture 5900 combines the two components shown in Figures 57 and 58. In this embodiment, the combined 6DoF pose estimate is achieved by combining AI-based keypoint detection and point cloud segmentation on both the NIR imagery from the LiDAR 5918 and the RGB imagery from the camera 5908. The transformation or mapping between the image spaces TNIR (T-nir) 5930 and TRGB (T-rgb) 5940 is achieved from a LUT or by calibration, respectively, as described above. In this embodiment, a heuristic or meta-model 5904 allows for determining how the various components of the 6DoF pose from various sources (5916, 5928) are combined (5926) to form the final 6DoF pose estimate. Depending on lighting conditions or other environmental information 5902, meta-model 5904 may use coordinates (x, y, z) of panel keypoints based on either camera image 5908 or LiDAR NIR image 5918 (or other LiDAR data channels, such as Lidar range 5932). For example, meta-model 5904 may configure the CV / AI pipeline to rely only on RGB camera 5908 for indoor operation, only on LiDAR reflectance data channel 6008 ( FIG. 60 ) for nighttime outdoor operation, and only on LiDAR NIR data channel 5918 for daytime outdoor operation. Furthermore, meta-model 5904 may appropriately weight both sets of keypoint coordinates or even optimize AI models (e.g., by selecting appropriate model weights), such as upper robot AI model 5912 for RGB and lower robot AI model 5938 for particular environmental conditions. Additionally, metamodel 5904 may invoke post-processing corrections 5924 by the lower robot CV / AI pipeline, as described below. According to some embodiments, pre-processing steps 5910 and 5920 and plane segmentation step 5934 are as described above.A database 5906 may store data from environmental sensors 5902 and provide the data to metamodel 5904 .

[0113] FIG. 60 is a schematic diagram of another sensor fusion architecture 6000 for determining the 6DoF pose of a solar panel, according to some embodiments. This example shows a sensor fusion architecture that fuses traditional analysis of LiDAR imagery 6008 and monocular AI 6004 based on images from an RGB camera 6002. Three possible methods for calculating the 6DoF pose are shown. One method uses a PnP solver 6006 with pixel locations of the four panel corners (x1, y1),..., (x4, y4) as input, and this method can be used indoors. Another method combines the world space coordinates (x, y, z) of panel keypoints with rotation angles (rx, ry, rz) based on a LUT mapping 6014 (from the LiDAR's NIR or reflectance image space to a point cloud). This method can be used outdoors. Yet another method is external calibration 6018 (aimed at mapping between the camera's image space and the point cloud). These various methods of calculating 6DoF pose may be combined in various ways (e.g., for indoor / outdoor use, etc.). Reflectance 6008, NIR 6010, and range 6012 may be used for night, day, and outdoor use, respectively. Although not shown, other LiDAR data channels, such as signal and reflectance, may be used in addition to or instead of the other data channels.

[0114] Figure 61 is a schematic diagram of a sensor fusion architecture 6100 according to some embodiments. Raw, uncorrected images from the NIR LiDAR 6102 are preprocessed 6104 using one of the techniques described above and processed by the upper robot AI 6106. First-stage post-processing 6108 provides a layer of redundancy. This first stage may include the following steps: [1] Determine corners from the NIR image (and / or other channels) in pixel space; [2] For each corner in pixel space, transform to world space coordinates and calculate the other corners given the panel dimensions; [3] For each set of four corners in world space, transform back to pixel space; [4] For the four sets of (four) corners in pixel space from [3] above, select a single corner based on some metric of convergence (e.g., Euclidean distance); [5] Transform the selected corner from (4) above in pixel space to world space and the set of (other corners) associated with it. [6] From the set of corners in [5] above, a reference corner (i.e., the left nearest) is used. In some embodiments, [7] if the variance of the pixel locations of the corners in [3] above exceeds a threshold, then downward robotic vision is utilized for a third estimate of 6DoF. The pixel locations of the selected corners are transformed (T) (6124) to obtain the (x, y, z) coordinates of the reference panel corner. This can be combined (6126) with the output of the planar and cylindrical segmentation 6122 (which generates output based on the point cloud output from the LiDAR distance 6120) to obtain a first estimate of the 6DoF (X, Y, Z, RX, RY, RZ) of the panel. In the second stage post-processing 6134, based on the first estimates, the locations of the four panel corners 6128 in coordinates (x, y, z) are further refined based on the two-pass homography algorithm described above. The locations 6128 are then transformed by the inverse transform T -1(6116) obtains the pixel locations 6110 of the four corners in image space. These pixel locations are input to a two-pass homography algorithm 6112 (an example of which is described above), which includes [8] determining a first-pass homography matrix H1, [9] determining a second-pass homography matrix H2, and

[10] back-projecting the four panel corners based on the inverse of H2. The output of step

[10] is input to a PNP solver 6114 to obtain a second estimate of the panel's 6DoF. This estimate can be input to a downward robot vision (AI) 6130 according to [7] above. Figure 61 clearly illustrates the redundancy layers provided by the first and second stage post-processing and the lower robot CV / AI pipeline vision. Although not shown in Figure 61, in some embodiments, this architecture can provide redundancy layers including RGB camera images, other LiDAR data channels such as signal and reflectance, and complementary AI models.

[0115] Range data channels from LiDARs (such as Ouster LiDARs) can be used for depth estimation as a type of segmentation, which is refined by "two-pass homography" post-processing before being input to the PnP solver. Additionally, depth estimation or segmentation may be performed using a stereo camera, whether in a single-baseline or multiple-baseline configuration, as described above.

[0116] Whether using a stereo camera or LiDAR, panel segmentation can be achieved by externally calibrating the camera to the robotic system (hand-eye). Calibration can also be performed by mapping the two image spaces using a homography transformation. For example, the image space of the LiDAR range image can be mapped to the image space of the camera in a calibration process that is purely image-based (i.e., does not involve robot motion).

[0117] Alternatively, the 6DoF pose of the panel may be determined solely by the LiDAR. The rotation angles (rx, ry, rz) of the panel can be determined by planar segmentation of point cloud data from the LiDAR. Furthermore, the (x, y, z) locations of keypoints (e.g., corners) of the panel may be determined by mapping keypoints found in other data channels (e.g., NIR, reflectance, etc.). To improve the determination of the latter keypoint locations (x, y, z), the LiDAR may be calibrated to a high-resolution camera (e.g., by extrinsic or homography, as described above) so that keypoints found in the camera image can be mapped or transformed to the LiDAR's image space.

[0118] Furthermore, the various approaches to 6DoF pose estimation (and associated post-processing) described above can be complemented by additional post-processing (of the initial 6DoF pose estimate) involving a different CV / AI pipeline, for example, using different sensors (e.g., cameras) attached to the lower robot. For example, the CV / AI pipeline for the lower robot may refine or correct the initial estimate of the panel's 6DoF pose by profiling the surface contours of the torque tube and / or clamps with structured light, such as laser line projection or grid line projection. The profiles of the torque tube, clamps, or other related structures, along with their known shapes and dimensions, can serve as the basis for refining the panel's initial 6DoF pose. Such refinement assumes that the panel installed on the torque tube is rotationally aligned along the tube and that the midpoint of the clamp coincides with the clamped edge of the installed panel.

[0119] 62 is a schematic diagram of a solar panel installation 6200 according to some embodiments. Cartesian axes X, Y, and Z are shown, as well as a panel reference 6202 for six degrees of freedom (6DoF) pose estimation.

[0120] 63 shows a flow diagram of a method 6300 of autonomous solar power installation according to some embodiments. The method includes acquiring 6302 one or more images during installation, the one or more images including images of one or more solar panels and an installation structure.

[0121] The method further includes preprocessing (6304) one or more images, including one or more of correcting for inherent camera characteristics or distortion, rectifying the image, and determining depth information. In some embodiments, this preprocessing includes correcting for camera distortion, rectifying the image, and / or determining depth information based on a single-baseline stereo camera, a multi-baseline stereo camera, a time-of-flight sensor, or a LiDAR sensor. Examples of preprocessing techniques are described above with reference to non-AI methods for pose estimation and exemplary AI models and inference according to some embodiments. The AI ​​model may be trained based on images with and without lens distortion correction.

[0122] The method further includes detecting 6306 the one or more solar panels by inputting the one or more images into one or more neural networks (in series or parallel) trained to detect solar panels. In some embodiments, the one or more neural networks are trained to output bounding boxes, segmentation, keypoints, depth, and / or 6DoF pose. These techniques are described above in the "Examples of AI Models and Inference" section according to some embodiments.

[0123] The method further includes a first post-processing (6308) for calculating a first panel pose based on the output of the one or more neural networks. This first post-processing may include a different AI / CV pipeline than that utilized for the pre-processing and / or neural networks. The goal may be to further refine the initial estimated panel pose.

[0124] In some embodiments, the first post-processing includes one or more computer vision algorithms (e.g., Ramer-Douglas-Peucker, Hough transform, Shi-Tomasi, homography transform, etc.) to determine the locations of panel keypoints by processing the output of one or more neural networks based on invariant structures in the image (e.g., invariant structures include grid lines and other visual references on the panels, shadows between panels, projected perspective lines, and the relative positions of the torque tube and the panels). Examples according to some embodiments are described above with reference to FIG. 51.

[0125] In some embodiments, the first post-processing further includes resolving perspective n-points based on the panel dimensions and the panel key points. In some embodiments, the panel key points are the four corners of the panel frame.

[0126] In some embodiments, the method further includes a second post-processing (6310) that includes one or more homography transformations (e.g., the two-pass homography algorithm described above) to obtain a second panel pose of the one or more solar panels based on the first panel pose. This second post-processing corrects or corrects inaccuracies in the first panel pose based on visual patterns or visual references on the solar panels. In some embodiments, these visual patterns comprise grid patterns on the solar panels.

[0127] In some embodiments, the output of the one or more neural networks includes panel segmentation. FIG. 52 illustrates an example application 5200 of segmentation according to some embodiments. FIG. 55 illustrates a schematic diagram of ROI segmentation 5500 according to some embodiments. This second post-processing includes obtaining masked panels by using the panel segmentation as a mask to filter out the background in the image (leaving only the foreground panels). The second post-processing further includes identifying grid intersections within the masked panels as corners by utilizing a corner-finding algorithm (e.g., Shi-Tomasi algorithm, Harris-Stephens algorithm, etc.). The second post-processing further includes calculating a homography matrix H1 using these corners. The second post-processing further includes determining the locations of the grid intersections in millimeter space based on H1. The second post-processing further includes converting the grid intersections in the image space to positions in millimeter space (a two-dimensional (2D) projection space in which the four (physical) corners of the solar panel have the following (x,y) coordinates: (0,0), (0,L), (W,0), (W,L), where W and L are the width and length of the solar panel, respectively) by calculating a homography matrix H2. The second post-processing further includes back-projecting the corners in millimeter space into image space using the inverse matrix of H2.

[0128] In some embodiments, the filtering comprises a binary AND between the mask and the raw image of the solar panel. In some embodiments, locating the grid intersections comprises matching in H1 space based on Euclidean distance (e.g., a portion of the shortest cell dimension, such as width).

[0129] In some embodiments, the second post-processing is performed in the image space k i The second post-processing step involves estimating the four corners of the solar panel at k iThe second post-processing step further involves calculating a homography matrix H1 that maps the pixel locations p of the panel features (as references) in image space based on H1. i and (ii) pixel position p in millimeter space. i Position q corresponding to i Further, the second post-processing includes identifying p i Aq i Further, the second post-processing step involves calculating the inverse matrix (H2 -1 ), involves backprojecting the corners (0,0), (H,0), (H,W), (0,W) of millimeter space into image space.

[0130] In some embodiments, p i and q i includes pixel structures other than the location of a grid intersection, the centroid of a diamond based on the grid intersection or the common vertex of two triangles, and / or the corners of a solar panel.

[0131] Various examples of utilizing homography transformation and sensor fusion architectures for 6DoF pose estimation of solar panels, according to some embodiments, are described above with reference to Figures 56-61. Specifically, the pre-processing and post-processing steps described above with reference to Figures 56-61 may be utilized in the pre-processing step, the first post-processing step, and / or the post-processing step of method 6300. For example, Figure 61 illustrates how a two-pass homography algorithm may be used for post-processing in any sensor fusion architecture based on corner detection, according to some embodiments.

[0132] In some embodiments, the installation structure includes a torque tube and a clamp, and the method further includes a third post-processing step including processing one or more images of the torque tube and / or the clamp. In some embodiments, the one or more images of the torque tube and / or the clamp are acquired using a high-resolution camera and structured illumination. In some embodiments, the structured illumination is a laser line that is approximately perpendicular or approximately parallel to the torque tube. In some embodiments, the processing of the one or more images of the torque tube and / or the clamp is performed by one or more neural networks and / or computer vision pipelines. In some embodiments, the method further includes locating a nut associated with the clamp using high-intensity illumination and a computer vision algorithm. In some embodiments, the high-intensity illumination is a ring light.

[0133] 63 , the method further includes generating 6312 control signals for operating the robotic controller to install the one or more solar panels based on the first panel pose. In some embodiments, the control signals are further generated 6314 based on the second panel pose. In various embodiments, the control signals are generated based solely on the first panel pose or solely on the second panel pose, or the control signals may be generated based on both the first panel pose and the second panel pose.

[0134] 35A shows a block diagram of an example image processing pipeline 3500 according to some embodiments. Pipeline 3500 includes a module 3502 for acquiring an image, a module 3504 for rectifying the image, a module 3506 for neural network image segmentation of the rectified image, a module 3508 for post-processing the output of module 3506 using computer vision techniques, a module 3510 for performing a Hough transform on the output of module 3508, a module 3512 for filtering and segmenting the Hough lines output by module 3510, a module 3514 for identifying intersections of the horizontal and vertical Hough lines output by module 3512, a module 3516 for estimating the pose of the panel based on the intersections of the horizontal and vertical Hough lines, and a module 3518 for publishing the pose estimate (e.g., using the 3D panel shape and the locations of corners in the image, etc.). Figure 35B shows an example corrected acquired image 3520 (output of modules 3502 and 3504) including an image of a solar panel 3522 and other objects 3524-2 (e.g., tape, etc.) and 3524-4 (e.g., wire, etc.). Figure 35C shows an example output 3526 (output of module 3506) for neural network image segmentation of the acquired corrected image shown in Figure 35B, according to some embodiments. Figure 35D shows an example panel corner detection 3528 (output of module 3514), according to some embodiments. In this example, corners 3530-2 and 3530-4 are detected based on horizontal lines 3532-4 and 3532-8 and vertical lines 3532-2 and 3532-6.

[0135] FIG. 36 shows example 3600 of road images 3602, 3606, and 3610 under various lighting conditions and image segmentation masks 3604, 3608, and 3612, according to some embodiments. Traditional computer vision techniques are useful when the environment is ideal. However, glare and overexposure / underexposure can have adverse effects on object detection algorithms. Machine learning techniques can overcome environmental inconsistencies, learn from examples with glare and lighting issues, and have the ability to generalize to new data during inference. Non-AI methods, such as those utilizing multi-baseline stereo, LiDAR / ToF, etc., may be used for non-ideal conditions. Furthermore, non-AI methods may be combined with AI-based methods.

[0136] Solar panel segmentation example Some embodiments perform solar panel segmentation by capturing images of the solar panel and torque tube under various lighting conditions. FIG. 37A shows an example of a captured image 3700 including a solar panel 3702 and a torque tube 3704, according to some embodiments. Some embodiments annotate the captured image of the solar panel. FIG. 37B shows an example of an annotated image 3706 (sometimes referred to as an annotated ground truth mask) of the captured image 3700, according to some embodiments. The annotated image includes a black background 3708, an outline of the torque tube 3712 shown in dark gray, and an outline of the solar panel 3710 shown in light gray. Some embodiments generate a dataset based on the annotated images, use the dataset to train an image segmentation model, and use the trained model to detect the solar panel and torque tube in poor lighting conditions. FIG. 37C shows an example prediction 3714 from the trained model, according to some embodiments. This trained model predicts the background 3708, the solar panel 3710, the torque tube 3712, and an object 3716 in the background (not shown in Figures 37A and 37B).

[0137] In some embodiments, images are continuously collected (and a dataset is built) and used to improve the accuracy of the model. Some embodiments improve the accuracy of the model by using human annotations. In some embodiments, the user can adjust the parameters of the segmentation model.

[0138] Some embodiments provide separate models for semantic segmentation and instance segmentation. Figure 38A illustrates an example of image classification. In this example, image classification detects the presence of a bottle 3802, a cup 3806, and a cube 3804. Figure 38B illustrates an example of object localization 3816 for the image shown in Figure 38A. In this example, rectangle 3808 locates bottle 3802, rectangle 3810 locates a first cube, rectangle 3812 locates cup 3806, and rectangles 3814-2 and 3814-4 locate cube 3804. Figure 38C illustrates an example of semantic segmentation 3818 according to some embodiments. Semantic segmentation helps identify label 3820 for bottle 3802, label 3822 for cube 3804, and label 3824 for cup 3806. FIG. 38D shows an example of instance segmentation 3826 according to some embodiments. Apart from identifying labels 3828 and 3832 for bottles 3802 and 3806, respectively, instance segmentation can determine labels 3830, 3834, and 3836 for cube 3804 to distinguish between instances of cube 3804. Instance segmentation can distinguish between multiple solar panel instances within a single image. Instance segmentation generates a mask for each class instance within a camera frame, enabling identification and localization of individual panels, using the same data collected for semantic segmentation and supporting lighting invariance. FIG. 39 shows an example of solar panel instance segmentation 3900 according to some embodiments. In this example, panel instances 3902, 3904, 3906, 3908, and 3910 are identified. This example shows instances of a panel (such as instances 3902 and 3904) each with a different orientation.

[0139] FIG. 40 illustrates an example image handling system 4000 according to some embodiments. The system 4000 includes multiple cameras: a camera 4002 for coarse positioning, a camera 4004 for capturing images as panels are picked, and a camera 4006 for capturing images as panels are placed. Camera 4002 includes a narrow field of view lens, while cameras 4004 and 4006 each include a wide field of view lens. Camera 4002 may be used to identify the trailer position and the robot initial position. In some embodiments, cameras 4004 and 4006 may be the same camera. Additionally, in some embodiments, camera 4002 may be used to locate clamps and a central structure during solar panel installation. Cameras 4002, 4004, and 4006 are coupled to image sensors 4008, 4010, and 4012 (e.g., AR0820 sensors), respectively. In some embodiments, these image sensors are optimized for both low light and / or high dynamic range performance. In some embodiments, system 4000 includes a high-speed digital video interface (e.g., FPD-Link, etc.) and Ethernet to connect the camera to one or more GPUs (e.g., a GPU 4014 suitable for edge AI processing such as Nvidia XT, a GPU suitable for image processing applications such as Nvidia AGX Xavier®, etc.). GPU 4016 implements the example image processing pipeline 3500 described above and is connected to robot controller 4018 using Ethernet. In some systems, GPU 4014 may be removed, and according to some embodiments, output from sensors may be connected directly to GPU 4016.

[0140] Some embodiments continue to capture training images while installing the solar panels. Figure 41 shows a trailer system 4100 with a coarse resolution camera 4102 that can be used to capture training images, according to some embodiments.

[0141] 42A and 42B show histograms 4200 and 4202 of the norm of the pose error of a neural network with and without a coarse position, according to some implementations. As shown, using a coarse position significantly reduces the error of the neural network (the difference between the actual position of the solar panel corners and the estimate from the neural network) (in some cases, down to about 5 inches to 0.7 inches).

[0142] 43A and 43B show examples 4300 and 4302 of approximate solar panel locations (e.g., locations 4304, 4306, 4308, and 4310) using the B Mask R-CNN model in accordance with some embodiments. Mask R-CNN is a convolutional neural network (CNN) used for image segmentation and instance segmentation. This deep neural network detects objects in an image and generates high-quality segmentation masks for each instance. Mask R-CNN is based on a region-based convolutional neural network. Image segmentation is the process of partitioning a digital image into multiple segments or sets of pixels that correspond to image objects. This segmentation is used to locate objects and boundaries (lines, curves, etc.). Mask R-CNN can be used for semantic segmentation and instance segmentation. Semantic segmentation classifies each pixel into a fixed set of categories without distinguishing between object instances. In other words, semantic segmentation corresponds to identifying / classifying similar objects as a single class at the pixel level. All objects are classified as a single entity (solar panel). Semantic segmentation is sometimes called background segmentation because it separates the subject matter of an image (e.g., solar panel, wires, etc.) from the background. On the other hand, instance segmentation (sometimes called instance recognition) addresses the accurate detection of all objects in an image and further provides accurate segmentation of each instance. In that sense, instance segmentation combines object detection, object localization, and object classification and helps identify instances of each object in an image. In Mask R-CNN, in addition to the two outputs for each candidate object, including the class label and bounding box offset, a third branch outputs the object mask.This mask output helps extract a finer spatial layout of the object. In addition to being easier to train, performing well, and being more efficient than other models, Mask R-CNN is particularly well-suited for solar panel identification because the neural network has the ability to perform both semantic and instance segmentation. Furthermore, the mask branch adds only a small amount of computational overhead, allowing for faster solar panel detection and more rapid experimentation. Mask R-CNN can be used for image segmentation, identifying objects in an image and generating masks within the object's boundaries.

[0143] 44 illustrates a system 4400 for solar panel installation, according to some embodiments. According to some embodiments, the system 4400 includes a main enclosure 4404, a battery enclosure 4402, an upper robotic end effector (EOAT) 4406, a lower robotic EOAT 4408, and a cradle 4410 for holding a solar panel 4414 on a trailer 4412.

[0144] Figure 45A shows a vision system 4502 mounted on a trailer and used to estimate the pose of the structure 4500, and Figure 45B shows a close-up view of the vision system 4502 according to some embodiments. Various embodiments may have the vision system mounted on different parts of the land vehicle, on a robotic arm, or at the end of an end effector.

[0145] FIG. 46A shows a vision system 4602 for module picking 4600, and FIG. 46B shows a close-up view of the vision system 4602 with a high-resolution camera with laser line generation according to some embodiments.

[0146] FIG. 47A shows a system 4700 for performing distance measurements at a module angle (i.e., when facing the module) between position 4702 (shown in an enlarged view in FIG. 47B) and position 4704 (shown in an enlarged view in FIG. 47C), according to some embodiments.

[0147] Figure 48A shows a laser line generation system 4800 for detecting the position of the tube and clamp according to some embodiments, Figure 48B shows a close-up view of the laser line generation system 4802, and Figure 48C shows a diagram 4804 of the laser line generation according to some embodiments (the horizontal line detects the clamp and the vertical line detects the tube).

[0148] Figure 49A shows a vision system 4900 for estimating the position of the tube and clamp and locating the nut on the clamp, according to some embodiments. Figure 49A also shows a socket wrench 4902 used to tighten the nut. Figure 49B shows a close-up of this vision system. As shown in Figure 49B, a camera uses the laser lines mentioned above to locate the tube and clamp, and a flash ring light to locate the nut on the clamp. These lasers allow for accurate estimation of the position of the tube and clamp. The flash ring light is used to locate the nut on the clamp, as shown in Figure 49C. When tightened, this nut compresses the clamp to hold the panel in place.

[0149] Example of solar panel installation using AI FIG. 50A shows a flow diagram of a method 5000 for autonomous solar power installation, according to some embodiments. The method includes acquiring 5002 images of the solar power installation in progress. The images include images of one or more solar panels and one or more torque tubes. In some embodiments, acquiring the images includes using one or more filters to avoid direct sunlight glare to detect the end effector (EOAT). In some embodiments, acquiring the images includes using a high-resolution camera with laser line generation to identify the location of one or more torque tubes and / or clamps. In some embodiments, the images include images of the clamps and / or central structure for the solar power installation in progress. In some embodiments, multiple images are acquired using a wide-angle fisheye lens to generate a composite HDR (high dynamic range) image within the camera hardware. These images are sent through the Robot Operating System (ROS), a high-level software framework for integrating robots and servos, and the images are corrected (e.g., from fisheye distortion to a flat image) using OpenCV (an image processing framework) modules. The region and bit depth are then selected and used to reduce the HDR image to a standard 8-bit image, which effectively crops the region and bit depth to adjust as input for the trained neural network. At the time of acquisition, the robot pose may be saved (using ROS) to generate a translation camera result relative to the trailer (trailer system used for solar panel installation). This may include the robot position and camera position to identify the location of the image in 3D space.

[0150] The method further includes detecting solar panel segments by inputting the images into a trained neural network trained to detect solar panels (5004). The neural network may be implemented using software and / or hardware (sometimes referred to as neural network hardware) using conventional CPUs, GPUs, ASICs, and / or FPGAs. In some embodiments, the trained neural network is composed of (i) a model for semantic segmentation to identify solar panel segments and (ii) a model for instance segmentation to identify multiple solar panels. In some embodiments, the trained neural network utilizes the Mask R-CNN framework for segmentation. These trained neural networks detect solar panel segments based on features extracted from images of ongoing solar installations. In some embodiments, the acquired images are input to the neural network via ROS (e.g., input images go from an OpenCV module to a neural network module). An example technique for training a neural network according to some embodiments is described below with reference to FIG. 55B. In some embodiments, the neural network performs image segmentation to identify the panel (or panels) without identifying the panel's location. In some embodiments, there is one model that performs both functions (semantic segmentation and instance segmentation). Some embodiments use two instances of the same model to optimize throughput. In such an example, the camera takes two images, one image passing through each instance. By running two models, twice as many images can be processed simultaneously.

[0151] The method further includes utilizing a computer vision pipeline to estimate 5006 a panel pose of one or more solar panels based on the solar panel segments. In some embodiments, the computer vision pipeline includes one or more computer vision algorithms for post-processing, Hough transform, filtering and segmenting Hough lines, finding horizontal and / or vertical Hough line intersections, and panel pose estimation using a given 3D panel shape and corner locations. In some embodiments, the computer vision pipeline estimates the panel pose by locating clamps and / or central structures. In some embodiments, the computer vision pipeline estimates the panel pose by locating one or more torque tubes and / or clamps. In some embodiments, the computer vision pipeline locates nuts. After locating the nuts, a socket wrench mounted on a smaller robotic arm engages and tightens the nuts to secure the panel in place. Prior to this step, the clamps are loose and the panel may fall due to wind.

[0152] In some embodiments, panel pose estimation is performed using conventional machine vision hardware to identify the panel's location in 3D space. In some embodiments, this is a rough identification of rounded edges and is not intended to be very precise. A Hough transform is then used to determine the exact location of the edges, followed by extrapolation of the panel's edge lines, determination of where the panels intersect, and identification of panel corners. The panel corners are published to identify the panel's location relative to the robot. For example, based on the panel's shape in 3D, the panel's pose is calculated based on the location of the panel corners in the image.

[0153] In some embodiments, to estimate the pose of the panel, the computer vision pipeline utilizes a PnP (Perspective n-Point) solver with the camera's intrinsic parameters (aware of the camera's inherent distortion and parallax). The extrinsic parameters then capture the camera's position relative to the robot using the pose of the robot arm and EOAT at the moment of image capture. The robot's pose may be continuously captured with a timestamp, which can then be used to match the robot's pose to the camera acquisition timestamp. In some embodiments, the computer vision pipeline utilizes the known pose of the robot arm and end effector (where the camera is mounted) at the time of image capture to calculate the position of one or more corners of the panel.

[0154] The method further includes generating 5008 control signals to operate a robot controller to install one or more solar panels based on the estimated panel pose. In some embodiments, after a panel is found, its location is projected along the tube, which locates clamp pixels to identify clamp locations (e.g., how far apart the clamp is, how close the clamp is to a clamp puller, etc.). In some embodiments, the clamp locations are used to verify that the clamps are located within a tolerance window required by the clamp puller on the EOAT. Some embodiments use a central structure to determine the sequence for placing one panel or two panels to avoid collisions with the fan gear. Some embodiments use the panel locations to verify that the trailer is in a valid position relative to the tube so that the robot is within reach of the tasks it needs to perform. Some embodiments use the pose of the leading panel to guide the lower robot's precise tube information acquisition, which drives the positions of the upper and lower robots for panel placement and nut driving. In some embodiments, the above-described precise tube information acquisition forms a profilometer system that uses horizontal and vertical lasers to locate the tube and clamp positions. This profilometer system refines the working position, which has a 10-20 mm error based on the coarse tube information, and reduces that error to less than plus or minus 5 mm. In the first panel, the error from the coarse tube is within 5 mm, but this becomes larger as it is projected, and is limited to less than plus or minus 5 mm by using the precise tube information.

[0155] FIG. 50B shows a flow diagram of a method 5010 for training a neural network for autonomous solar power installations, according to some embodiments. The method includes acquiring multiple images of a solar panel installation under various lighting conditions (5012), annotating the multiple images to identify solar panel images (5014) (human-annotated images may be used instead of or in addition to automatically annotated images), and detecting solar panels in poor lighting conditions by training one or more image segmentation models using the solar panel images (5016). In some embodiments, the neural network is manually trained using a variety of images, such as multiple images of clamps and panels on different backgrounds (e.g., grass, dirt, etc.), with different panel counts, and under different weather conditions (e.g., sunny conditions, rainy conditions, etc.). Lines are drawn within these images to indicate which pixels represent panels, clamps, tubes, and central structures. Using these images and their masks, a series of pseudo-images (sometimes called image augmentation) are generated, which are then used by the neural network in the training process. These pseudo-images are input images with angular distortions to allow multiple training runs using the same input image. For example, 300-1000 real (input) images may be used for training, and 10-20 pseudo-images may be generated for each real image.

[0156] Some embodiments of the present invention have been described above with the aid of functional building blocks illustrating the implementation of specific functions and relationships thereof. The boundaries of these functional building blocks are arbitrarily defined herein for the convenience of description. Alternative boundaries can be defined so long as the specific functions and relationships thereof are appropriately implemented.

[0157] It will be apparent to those skilled in the art that various modifications and variations can be made to the system for installing solar panels of the present invention without departing from the spirit or scope of the present invention. Therefore, it is intended that the present invention encompasses the modifications and variations of the present invention within the scope of the appended claims and their equivalents. It should be understood that the phrases or terms used herein are for the purpose of description and not limitation, and should therefore be interpreted by those skilled in the art in light of the teachings and guidance.

[0158] The breadth and scope of the present invention should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the following claims and their equivalents. [Explanation of symbols]

[0159] 100 Arm Assembly Tools, Robot Arms 102 frames 102-A Truss 104 Wearable Devices 106 Linear guide assembly 108 Clamping tool, linear movable clamping tool 108-A Engagement member 108-B Proximity Sensor 110 Force torque transducer, force torque actuator 112 Junction box, controller 120 solar panels 602 Clamp Assembly 604 Torque tube 604 Installation structure 606 Roller 608 Spring Mechanism 610 Receptor Channel 802 Optical Sensor 802-A position 802-B position 804 Orientation Assembly 903 Assembly Mobile Robot 905 Storage Container 907 Land Vehicles 908 Upper Section 909 Lower Section 911 Second Robot Arm 1005 Module Vehicle 3500 Image Processing Pipeline 3502 Module 3504 Module 3506 Module 3508 Module 3510 Module 3512 Module 3514 Module 3516 Module 3518 Module 3520 Corrected acquired image 3522 Solar Panels 3524-2 Object 3524-4 Object 3526 Output 3528 Panel corner detection 3530 Corner 3532-4 Horizontal line 3532-8 Horizontal line 3532-2 vertical line 3532-6 Vertical Line 3600 Example images 3602, 3606, and 3610 of a road under various lighting conditions and image segmentation masks 3604, 3608, and 3612 3602 Road Images 3604 Segmentation Mask 3606 Road Images 3608 Segmentation Mask 3610 Road Images 3612 Segmentation Mask 3700 Captured Images 3702 Solar Panels 3704 Torque tube 3706 annotated images 3708 Black Background 3710 Solar Panels 3712 Torque tube 3714 Predicting a single example using a trained model 3716 objects 3802 bottles 3804 Cube 3806 cups 3808 rectangle 3810 rectangle 3812 rectangle 3814 rectangle 3816 Object Localization 3818 Semantic Segmentation 3820 Label 3822 Label 3824 Label 3826 Instance Segmentation 3828 Label 3830 Label 3900 Segmentation 3902 Panel Instance 4000 Image Handling System 4002 Camera 4004 Camera 4006 Camera 4008 Image Sensor 4018 Robot Controller 4100 Trailer System 4102 Coarse Resolution Camera 4200 Histogram 4300 Example of approximate solar panel locations (e.g., locations 4304, 4306, 4308, and 4310) using the B Mask R-CNN model. 4304 Approximate location of solar panels 4306 Approximate location of solar panels 4308 Approximate location of solar panels 4310 Approximate location of solar panels 4400 System 4402 Battery Enclosure 4404 Main Enclosure 4406 Upper Robot End Effector (EOAT) 4408 Downward Robot EOAT 4410 Cradle 4412 Trailer 4414 Solar Panels 4500 Structures 4502 Vision System 4600 Module Picking 4602 Vision System 4700 System 4702 position 4704 position 4800 Laser Line Generator System 4802 Laser Line Generator System 4900 Vision System 4902 Socket wrench 5100 Applications 5102 Hough Line 5104 Hough Line 5110 Solar Panels 5200 apply 5300 Example 5400 detection 5402 grid intersections 5500 ROI Segmentation 5502 12MP RGB Camera 5504 12MP RGB camera 5506 AI Pipeline 5600 Multi-Data Channel Performance 5602 distance 5604 Signal 5606 Ambient light 5608 Reflectance 5700 Computer Vision / AI Architecture, CV / AI Pipeline 5702 Ambient NIR Grayscale Image 5704 Conversion 5706 Point Cloud Distance Data 5708 Pretreatment, pretreatment step 5710 AI 5712 Plane Segmentation, Plane Segmentation Step 5714 Conversion or Mapping 5716 Combination 5800 Computer Vision / AI Pipeline 5802 Camera 5806 LiDAR 5808 Pre-processing step 5810 AI Steps 5812 Plane Segmentation Step 5816 Combination 5900 Architecture 5900 Sensor Fusion Architecture 5902 Lighting conditions or other environmental information, environmental sensors 5904 Metamodel 5906 Database 5908 camera, RGB camera, camera image 5910 Pretreatment Step 5912 Upper Robot AI Model 5916 Source 5918 LiDAR NIR images, LiDAR NIR data channels 5924 Post-processing correction 5926 Combinations of various components of 6DoF posture 5928 Source 5930 Image Space TNIR (T-nir) 5932 Lidar distance 5934 Planar Segmentation Step 5938 Downward Robot AI Model 5940 Image Space TRGB (T-rgb) 6000 Sensor Fusion Architecture 6002 RGB camera 6004 Monocular AI 6006 PnP Solver 6008 LiDAR reflectance data channels, LiDAR imagery, reflectance 6010 NIR 6012 distance 6014 LUT mapping 6018 External Calibration 6100 Sensor Fusion Architecture 6102 NIR LiDAR 6104 Pretreatment 6106 Upper Robot AI 6108 Post-processing 6110 pixel position 6112 Two-Pass Homography Algorithm 6114 PNP solver 6120 LiDAR distance 6122 Planar and Cylinder Segmentation 6124 Conversion (T) 6126 Combination 6128 Panel corner, position 6130 Downward Robot Vision (AI) 6134 Post-processing 6200 Solar panel installation 6202 Panel Standards 6308 First post-processing 6310 Second post-processing

Claims

1. 1. A method for autonomous solar panel installation, comprising: acquiring one or more images during installation, the one or more images including images of the one or more solar panels and an installation structure; pre-processing the one or more images, including one or more of correcting for inherent camera characteristics or distortions, rectifying the image, and determining depth information; detecting the one or more solar panels by inputting the one or more images into one or more neural networks trained to detect solar panels; performing a first post-processing step to calculate a first panel pose based on the output of the one or more neural networks; generating a control signal for operating a robotic controller for installing the one or more solar panels based on the first panel attitude; A method comprising:

2. performing second post-processing based on the first panel pose, the second post-processing including one or more homography transformations to obtain second panel poses of the one or more solar panels; further comprising The second post-processing step corrects or corrects the inaccuracy of the first panel attitude based on a visual pattern or visual reference on the solar panel; 10. The method of claim 1, wherein the control signals for operating the robotic controller to install the one or more solar panels are further based on the second panel pose.

3. The method of claim 2 , wherein the visual pattern comprises a grid pattern on the solar panel.

4. the output of the one or more neural networks comprises a panel segmentation; The step of performing second post-processing includes: filtering out the background of the image by using the panel segmentation as a mask to obtain a masked panel; using a corner-finding algorithm to identify grid intersections within the masked panel as corners; Using the corners, the homography matrix H 1 and calculating H 1 determining the locations of the grid intersections in millimeter space based on A homography matrix H is used to transform the grid intersections in image space to their positions in millimeter space. 2 and calculating H 2 back-projecting the corners in the millimeter space into image space using the inverse matrix of 3. The method of claim 2, comprising:

5. The step of determining the locations of the grid intersections is performed using a Euclidean distance-based H 1 The method of claim 4, including matching in space.

6. the output of the one or more neural networks comprises a panel segmentation; The step of performing second post-processing includes: filtering out the background of the image by using the panel segmentation as a mask to obtain a masked panel; identifying grid intersections within the masked panel in millimeter space using one or more artificial intelligence techniques; Based on the grid intersections, a homography matrix H 1 and calculating H 1 back-projecting the corners in the millimeter space into image space using the inverse matrix of 3. The method of claim 2, comprising:

7. The step of performing second post-processing includes: The four corners of the solar panel in the image space, k i and estimating k i The homography matrix H that maps 1 and calculating H 1 Based on this, (i) the pixel location p of the panel feature in image space i and (ii) the pixel position p in millimeter space. i Position q corresponding to i and identifying p i Aq i The homography matrix H that maps to 2 and calculating H 2 -1 backprojecting the corners (0,0), (H,0), (H,W), (0,W) from millimeter space to image space based on 3. The method of claim 2, comprising:

8. 8. The method of claim 1, wherein the one or more neural networks are trained to output bounding boxes, segmentations, keypoints, depth, and / or 6DoF poses.

9. 9. The method of claim 1, wherein the pre-processing step comprises correcting camera distortion, rectifying the images, and / or determining depth information based on a single-baseline stereo camera, a multi-baseline stereo camera, a time-of-flight sensor, or a LiDAR sensor.

10. 10. The method of claim 1, wherein the first post-processing step comprises one or more computer vision algorithms for determining positions of panel key points by processing the output of the one or more neural networks based on invariant structures in the image.

11. The method of claim 1 , wherein the step of performing a first post-processing further comprises resolving perspective n-points based on panel dimensions and panel key points.

12. The method of claim 11 , wherein the panel key points are the four corners of the panel frame.

13. the mounting structure includes a torque tube and a clamp; 13. The method of claim 1, further comprising a third post-processing step comprising processing one or more images of the torque tube and / or the clamp.

14. 14. The method of claim 13, wherein the one or more images of the torque tube and / or the clamp are acquired using a high-resolution camera and structured illumination.

15. 15. The method of claim 14, wherein the structured illumination is a laser line that is approximately perpendicular or approximately parallel to the torque tube.

16. 14. The method of claim 13, wherein the processing of the one or more images of the torque tube and / or the clamp is performed by one or more neural networks and / or computer vision pipelines.

17. 17. The method of claim 13, further comprising locating a nut associated with the clamp by utilizing high intensity illumination and a computer vision algorithm.

18. The method of claim 17 , wherein the high intensity light is a ring light.

19. A system for installing solar panels, a camera system for capturing one or more images during installation, the one or more images including images of the one or more solar panels and an installation structure; One or more devices for (i) estimating a panel pose of the one or more solar panels based on the solar panel segments by pre-processing the one or more images, including one or more of correcting inherent camera characteristics or distortions, rectifying the images, and determining depth information; (ii) detecting the one or more solar panels based on the one or more images; and (iii) performing first post-processing to calculate a first panel pose based on the output of the one or more neural networks; a controller for generating a control signal for operating a robotic controller for installing the one or more solar panels based on the first panel attitude; A system comprising:

20. the one or more devices for second post-processing, including one or more homography transformations to obtain second panel poses of the one or more solar panels based on the first panel pose. Furthermore, the second post-processing compensates for or corrects inaccuracies in the first panel pose based on a visual pattern or visual reference on the solar panel; 20. The system of claim 19, wherein the control signal for operating the robotic controller to install the one or more solar panels is further based on the second panel pose.