Autonomous solar installation using artificial intelligence

By using arm-end assembly tools and machine learning technology in the solar panel processing system, the problem of low installation efficiency and reliability of solar panels in the prior art is solved, an efficient and reliable installation process is achieved, and costs are reduced.

CN120051932APending Publication Date: 2025-05-27THE AES CORPORATION
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202380068382.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-08-11
Filing Date
2023-08-11
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently and reliably install solar panels, especially in rotatable structures, and are costly to install.

Method used

A solar panel processing system is used that includes arm-end assembly tools and machine learning technology. The arm end assembly tool consists of a frame, suction cup and linear guide assembly, which is attached to the solar panel using the suction cup. The linear guide assembly is precisely positioned and fixed through a clamping tool and a force torque sensor. Machine learning techniques are used to deal with environmental inconsistencies, generalizing to new data by learning to deal with glare and lighting problems in examples.

Benefits of technology

Improves the installation efficiency and reliability of solar panels, reduces installation costs, and maintains efficient performance under different environmental conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120051932A_ABST
    Figure CN120051932A_ABST
Patent Text Reader

Abstract

Systems and methods for mounting a solar panel are provided. The method includes obtaining an image of the solar panel and the mounting structure during installation. The method further includes pre-processing the image by compensating for camera internal parameters or distortions, correcting the image, and / or determining depth information. The method further includes detecting the solar panel by inputting the image into a neural network. The method also includes a first post-processing to calculate a first panel pose based on the output of the neural network. The method further includes generating a control signal based on the first panel pose for operating a robot controller to install the solar panel. In some embodiments, the method further includes a homographic transformation to obtain a second panel pose based on the first panel pose and a visual pattern or reference on the solar panel, and further to generate a control signal based on the second panel pose.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Related Application Data

[0002] This application claims priority to U.S. Provisional Application No. 63 / 397,125, filed on Aug. 11, 2022, and is hereby incorporated by reference in its entirety under 35 U.S.C. § 119. Technical Field

[0003] The present disclosure generally relates to solar panel handling systems, and more particularly to systems and methods for mounting solar panels on mounting structures. Background Art

[0004] In the following discussion, reference is made to certain structures and / or methods. However, the following reference should not be construed as an admission that such structures and / or methods constitute prior art. Applicants expressly reserve the right to demonstrate that such structures and / or methods do not qualify as prior art against the present invention.

[0005] The installation of a photovoltaic array generally involves securing solar panels to a mounting structure. This bottom support provides attachment points for the individual solar panels, as well as facilitates the routing of the electrical system and (if applicable) any mechanical components. Due to the fragile nature and large size of solar panels, the process of securing solar panels to a mounting structure poses unique challenges. For example, in many cases, the solar panels of a photovoltaic array are mounted on a rotatable structure that can rotate the solar panels about an axis to enable the array to track the sun. In such cases, it is difficult to ensure that all of the solar panels in the array are coplanar and level with respect to the axis of the rotatable structure. Additionally, the installation cost of a photovoltaic array can represent a significant portion of the total construction cost of the photovoltaic array. Accordingly, there is a need for a more efficient and reliable solar panel handling system for installing solar panels in a photovoltaic array. When the environment is ideal, conventional computer vision techniques can be used.

[0006] However, glare, overexposure, or underexposure can negatively affect object detection algorithms. Summary of the Invention

[0007] Accordingly, the present invention relates to a solar panel handling system that substantially eliminates one or more problems due to the limitations and disadvantages of the related art.

[0008] The solar panel handling system disclosed herein facilitates the installation of solar panels of a photovoltaic array on a pre-existing mounting structure (e.g., a torque tube). By combining tools for handling solar panels with components enabling the solar panels to match a solar panel support structure, the installation of solar panels can be made more efficient and reliable. Some embodiments use machine learning techniques to overcome environmental inconsistencies. The system can learn from examples with glare and lighting issues and can generalize to new data during inference.

[0009] Additional features and advantages of the invention will be set forth in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. The objectives and other advantages of the invention will be realized and attained by means of the instrumentalities and combinations particularly pointed out in the written description and claims hereof as well as the appended drawings.

[0010] In accordance with the purposes of the present invention, as embodied and broadly described, a system for installing a solar panel may include an end-of-arm assembly tool that includes a frame and a suction cup coupled to the frame, and a linear guidance assembly coupled to the end-of-arm assembly tool, wherein the linear guidance assembly includes: a linearly movable clamping tool that includes an engagement member configured to engage a fixture assembly slidably coupled to a mounting structure; a force-torque sensor configured to move the clamping tool along the mounting structure; and a junction box coupled to the frame and including a controller configured to control the force-torque sensor and the suction cup, and a power source.

[0011] In another aspect, a method of installing a solar panel may include: engaging an end-of-arm assembly tool with the solar panel, the end-of-arm assembly tool including a frame and a suction cup coupled to the frame; positioning the solar panel relative to a mounting structure that has a fixture assembly slidably coupled thereto; engaging a linear guidance assembly coupled to the end-of-arm assembly tool with the fixture assembly, the linear guidance assembly including a linearly movable clamping tool that includes an engagement member configured to engage the fixture assembly and a moment sensor configured to move the clamping tool along the mounting structure; and actuating the force-torque sensor to move the fixture assembly along the mounting structure to engage one side of the solar panel, thereby fixing the solar panel relative to the mounting structure.

[0012] In another aspect, according to some embodiments, a method for training a neural network for autonomous solar installation is provided. The method includes obtaining one or more images during installation. The one or more images include images of one or more solar panels and the installation structure. The method further includes preprocessing the one or more images, including compensating for camera internal parameters or distortion, rectifying the images, and determining depth information, among other things. The method further includes detecting the one or more solar panels by inputting the one or more images into one or more neural networks trained to detect solar panels. The method further includes a first post-processing to calculate a first panel pose based on the outputs of the one or more neural networks. The method further includes generating a control signal based on the first panel pose for operating a robotic controller to install the one or more solar panels. In some embodiments, the method further includes a second post-processing, which includes one or more homography transformations to obtain a second panel pose of the one or more solar panels based on the first panel pose. The second post-processing compensates for or corrects inaccuracies in the first panel pose based on visual patterns or fiducials on the solar panels. The control signal for operating the robotic controller to install the one or more solar panels is also based on the second panel pose.

[0013] It is to be understood that both the foregoing general description and the following detailed description are exemplary and explanatory and are intended to provide further explanation of the claimed invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate the invention and, together with the description, further serve to explain the principles of the invention and enable a person skilled in the relevant art to make and use the invention. The exemplary embodiments are best understood when read in conjunction with the accompanying drawings. It should be emphasized that, by convention, the features of the drawings are not drawn to scale. Instead, for clarity, the dimensions of various features are arbitrarily enlarged or reduced.

[0015] The accompanying drawings include the following figures:

[0016] Figure 1 A perspective view of a solar panel handling system and a container for solar panels, according to an embodiment of the present disclosure, is shown.

[0017] Figures 2A - 2C Correspondingly shown is Figure 1 a top view, a front view, and a side view of the solar panel handling system and the container for solar panels.

[0018] Figure 3A Correspondingly shown are a top view, a front view, and a side view of a solar panel handling system coupled to a single solar panel, according to an embodiment of the present disclosure.

[0019] Figure 4A and 4B shows a perspective view of a solar panel processing system according to an embodiment of the present disclosure.

[0020] Figure 5A and 5B correspondingly shows a top view and a front view of a solar panel processing system according to an embodiment of the present disclosure.

[0021] Figure 5C shows a side view of a clamping tool of a solar panel processing system according to an embodiment of the present disclosure in a retracted position.

[0022] Figure 5D shows a side view of a clamping tool of a solar panel processing system according to an embodiment of the present disclosure in an extended or advanced position.

[0023] Figure 6A and 6B shows a perspective view of a clamping tool of a solar panel processing system according to an embodiment of the present disclosure engaged with a fixture assembly coupled to a mounting structure.

[0024] Figure 7A shows a top view of a clamping tool of a solar panel processing system according to an embodiment of the present disclosure engaged with a fixture assembly coupled to a mounting structure.

[0025] Figure 7B shows a front view of a clamping tool of a solar panel processing system according to an embodiment of the present disclosure engaged with a fixture assembly coupled to a mounting structure.

[0026] Figure 7C shows a side view of a clamping tool of a solar panel processing system according to an embodiment of the present disclosure engaged with a fixture assembly coupled to a mounting structure.

[0027] Figure 7D shows a rear view of a clamping tool of a solar panel processing system according to an embodiment of the present disclosure engaged with a fixture assembly coupled to a mounting structure.

[0028] Figure 8 schematically shows a solar panel processing system according to an embodiment of the present disclosure during the process of installing a solar panel in a top view.

[0029] Figure 9 illustrates a solar panel processing system that includes an assembly tool coupled to an assembly mobile robot using a robotic arm.

[0030] Figure 10 illustrates a solar panel processing system having two robotic arms, wherein two assembly tools are coupled to an assembly mobile robot using respective robotic arms.

[0031] Figures 11A to 11C Illustrates a process for installing a solar panel.

[0032] Figure 12A And 12B Illustrates an arrangement of a mobile robot system for a modular vehicle including two modules and a ground vehicle having two robotic arms.

[0033] Figure 13 Schematically illustrates an installation implemented using computer vision registration.

[0034] Figure 14 Schematically illustrates an arrangement in which the modular vehicle is replaced with a new modular vehicle having additional solar panels for supplementation.

[0035] Figures 15 - 34 Provides a detailed illustration of an exemplary construction of a system for installing a solar panel according to an embodiment of the present disclosure.

[0036] Figure 35A Shows a block diagram of an exemplary image processing pipeline according to some embodiments.

[0037] Figure 35B Shows an exemplary rectified acquired image according to some embodiments.

[0038] Figure 35C Shows according to some embodiments for Figure 35B An exemplary output of neural network image segmentation for acquiring a rectified image as shown in.

[0039] Figure 35D Shows an exemplary panel corner detection according to some embodiments.

[0040] Figure 36 Shows examples of road images and segmentation masks of the images under different lighting conditions according to some embodiments.

[0041] Figure 37A Shows an example of a captured image including a solar panel and a torque tube according to some embodiments.

[0042] Figure 37B Shows according to some embodiments Figure 37A An example of an annotated image of the captured image as shown in.

[0043] Figure 38A Shows an example of image classification.

[0044] Figure 38B Shows Figure 38A An example of object localization of the image as shown in.

[0045] Figure 38C Shows an example of semantic segmentation according to some embodiments.

[0046] Figure 39 Shows an example of instance segmentation of a solar panel according to some embodiments.

[0047] Figure 40 Shows an exemplary image processing system according to some embodiments.

[0048] Figure 41 shows a trailer system according to some embodiments.

[0049] Figure 42A and 42B Shows a histogram of the pose error norm of a neural network with and without a rough position according to some embodiments.

[0050] Figure 43A and 43B Shows an example of the rough position of a solar panel using a B Mask R-CNN model according to some embodiments.

[0051] Figure 44 Shows a system 4400 for solar panel installation according to some embodiments.

[0052] According to some embodiments, Figure 45A Shows a vision system for tracking the trailer position, and Figure 45B Shows a magnified view of the vision system.

[0053] According to some embodiments, Figure 46A Shows a vision system for module picking, and Figure 46B Shows a magnified view of the vision system.

[0054] Figure 47A 、 47B and 47C show a system 4700 for distance measurement at a module angle according to some embodiments.

[0055] Figure 48A Shows a system for laser line generation for detecting the positions of tubes and fixtures according to some embodiments.

[0056] According to some embodiments, Figure 48B Shows Figure 48A a magnified view of the laser line generation system shown in Figure 48C and shows a view of laser line generation (the horizontal line detects the fixture and the vertical line detects the tube).

[0057] According to some embodiments, Figure 49AShows a vision system 4900 for estimating the position of tubes and fixtures, and Figure 49B shows an enlarged view of the vision system.

[0058] Figure 50A Shows a flowchart of a method for autonomous solar installation according to some embodiments.

[0059] Figure 50B Shows a flowchart of a method for training a neural network for autonomous solar installation according to some embodiments.

[0060] Figure 51 Shows an exemplary application of the Hough transform for identifying the corners of a solar panel according to some embodiments.

[0061] Figure 52 Shows an exemplary application of segmentation according to some embodiments.

[0062] Figure 53 Shows an exemplary application of a corner detection algorithm according to some embodiments.

[0063] Figure 54 Shows grid intersections that can be detected using a two-pass homography algorithm according to some embodiments.

[0064] Figure 55 Shows a schematic diagram of region of interest segmentation according to some embodiments.

[0065] Figure 56 Shows the multi-data channel capabilities of LiDAR used in some embodiments.

[0066] Figure 57 Is a schematic diagram of an exemplary computer vision / artificial intelligence (AI) architecture for estimating the six degrees of freedom (6DoF) pose of a solar panel according to some embodiments.

[0067] Figure 58 Is a schematic diagram of another exemplary computer vision / AI pipeline for estimating the 6DoF pose of a solar panel according to some embodiments.

[0068] Figure 59 Is a schematic diagram of a sensor fusion architecture for determining the 6DoF pose of a solar panel according to some embodiments.

[0069] Figure 60 Is a schematic diagram of another sensor fusion architecture for determining the 6DoF pose of a solar panel according to some embodiments.

[0070] Figure 61 Is a schematic diagram of a sensor fusion architecture according to some embodiments.

[0071] Figure 62 It is a schematic diagram of the installation of a solar panel according to some embodiments.

[0072] Figure 63 It shows a flowchart of a method for autonomous solar installation according to some embodiments.

[0073] The features and advantages of the present invention will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings, in which like reference numerals always denote corresponding elements. In the drawings, like reference numerals generally represent identical, functionally similar, and / or structurally similar elements. Detailed Description

[0074] Reference will now be made in detail to embodiments of the present invention, which are illustrated in the accompanying drawings.

[0075] Figure 1 It shows a perspective view of a solar panel handling system and a case of solar panels according to an embodiment of the present disclosure. The solar panel handling system may include an end-of-arm tooling 100, which may be coupled to respective solar panels 120 from a case of solar panels and move them to positions relative to a mounting structure for installation.

[0076] The end-of-arm tooling 100 may include a frame 102 and one or more attachment devices 104 coupled to the frame 102. Exemplary attachment devices 104 include suction cups or other structures that may be releasably attached to the surface of the solar panel 120 and, at least generally, maintain the attachment during manipulation of the solar panel 120 by the end-of-arm tooling 100. The frame 102 may be composed of a plurality of trusses 102-A for providing structural strength and stability to the frame 102. The frame 102 also serves as a base for the end-of-arm tooling 100 and other related components of the solar panel handling system disclosed herein.

[0077] Other related components of the solar panel handling system disclosed herein may be coupled to the frame 102 to fix the relative positions of the components on the end-of-arm tooling 100. One or more of the various components of the solar panel handling system may be coupled to one or more of the trusses 102-A to fix the relative positions of the components on the end-of-arm tooling 100.

[0078] For example, the attachment device 104 is configured to reliably attach to a planar surface, such as by using vacuum to attach to a surface such as a solar panel. In a suction cup embodiment, the suction cup can be actuated by pushing the disk against the planar surface, thereby pushing air out of the disk and forming a vacuum seal with the planar surface. Thus, the planar surface adheres to the suction cup with a certain adhesion strength, which depends on the size of the suction cup and the integrity of the seal with the planar surface. In some embodiments, the suction cup engages with the solar panel to form an airtight seal, and then a vacuum pump sucks air out of the suction cup, thereby creating the vacuum required for proper adhesion to the solar panel. In some embodiments, when the planar surface is sealed to the suction cup, an air inlet (not shown) supplies air onto the planar surface in order to deactivate the vacuum and release the planar surface from the suction cup.

[0079] The system may also include a linear guiding assembly 106 coupled to the end effector tool 100. The linear guiding assembly 106 includes a linearly movable gripping tool 108 having an engagement member 108-A that is configured to engage a fixture assembly coupled to a mounting structure. The linear guiding assembly 106 can be actuated to move the gripping tool 108 along an axis between, for example, an extended position and a retracted position. The movement axis of the gripping tool 108 can be parallel to the axis of the mounting structure. Thus, the linear guiding assembly 106 can move the gripping tool 108 and the engagement member 108-A along the mounting structure.

[0080] In some embodiments, the engagement member 108-A can include an electromagnet that can be actuated to grip the fixture assembly 602 (see Figure 6A , 6B ). Alternatively or additionally, the engagement member 108-A can include a gripper to prevent detachment between the fixture assembly 602 and the engagement member 108-A when the linear guiding assembly 106 is actuated to move the gripping tool relative to the mounting structure, as described in more detail elsewhere herein.

[0081] The linear guiding assembly 106 is actuated using a force torque sensor 110. In some embodiments, the linear guiding assembly 106 and the force torque sensor 110 can form a rack and pinion structure such that rotation of the force torque sensor 110 causes advancement or retraction of the gripping tool 108. In some embodiments, the linear guiding assembly 106 can be a hydraulic assembly that includes a telescopic shaft coupled to the gripping tool 108. In such an embodiment, the force torque sensor 110 can be configured in the form of a pump for pumping hydraulic fluid. In other embodiments, the force torque sensor 110 can be configured in the form of a linear drive motor or coupled to a linear drive motor that engages a surface of a telescopic shaft coupled to the gripping tool 108.

[0082] In some embodiments, the linear guidance assembly 106 may include an electric rod actuator to move the gripping tool 108 parallel to the axis of the mounting structure.

[0083] In some embodiments, the guidance assembly 106 may include rollers 606 to assist the gripping tool 108 in moving along the mounting structure 604. For example, the roller may include a bearing or other components designed to reduce friction when the gripping tool 108 moves relative to the mounting structure. The roller may be coupled to a sensor, such as a force sensor or a rotational sensor, to provide feedback to the controller.

[0084] In some embodiments, the guidance assembly may include a spring mechanism 608 that enables the gripping tool 108 to tilt slightly (up to 15 degrees) relative to the mounting structure 604. This tilt may occur when the orientation assembly 804 tilts the arm-end assembly tool 100 relative to the mounting structure 604 to properly level the solar panel.

[0085] The system may also include a junction box 112 coupled to the frame 102. The junction box 112 may include a controller configured to control the force torque sensor 110 and the attachment device 104. In some embodiments, the junction box 112 may also include a power source or a power controller for controlling the power supply to the various components.

[0086] In some embodiments, the controller 112 may include a processor operatively coupled to a memory. The controller 112 may receive inputs from sensors associated with the solar panel processing system (e.g., the optical sensors or proximity sensors 108-B described elsewhere herein). The controller 112 may then process the received signals and output control commands for controlling one or more components (e.g., the linear guidance assembly 106, the gripping tool 108, or the attachment device 104). For example, in some embodiments, the controller 112 may receive a signal from a proximity sensor that determines that the fixture assembly is approaching the trailing edge of the solar panel being installed and may accordingly reduce the speed of the linear guidance assembly 106 to reduce excessive force and impact on the solar panel.

[0087] Reference Figure 8 , in some embodiments, the solar panel processing system may further include an optical sensor 802, such as a camera, a photodetector, or any other optical imaging or light-sensing device. The optical sensor is appropriately located on the frame 102, e.g., at the outer surface or the lower surface of the edge member shown in position 802-A in Figure 8 or at an internal position of the frame 102 that has a field of view including the leading edge of the solar panel, such as Figure 8as shown at position 802-B therein. The optical sensor can be configured to sense the orientation of the solar panel relative to the mounting structure during operation of the end-of-arm assembly tool. In some embodiments, the optical sensor can be configured in the form of one or more light guided levels (not shown). In such embodiments, one or more light beams (e.g., laser beams) can be projected from one end of the end-of-arm assembly tool 100 (e.g., a first position on the frame 102) along or parallel to the axis of the mounting structure 604. One or more photodetectors can be positioned at the other end of the end-of-arm assembly tool 100, e.g., at a second position on the frame 102, to detect the one or more laser beams. Thus, if the solar panel 120 being installed is not properly oriented or leveled relative to the mounting structure 604, the solar panel 102 may block some or all of the one or more laser beams, resulting in a varying signal from the one or more photodetectors, thereby indicating that the solar panel 120 is not properly oriented or leveled relative to the mounting structure 604.

[0088] In some embodiments, one or more sensors (e.g., the optical sensor 802) can be used to detect and identify objects in order to position and control the installation with improved accuracy. The sensors can be implemented in conjunction with a neural network of, for example, an artificial intelligence (AI) system. For example, the neural network can include acquiring and correcting images related to the solar panel processing system, solar panels (both installed and to be installed), and the installation environment (both the natural environment, e.g., terrain; and installed equipment, e.g., structures associated with the solar panel array). Additionally, for example, the neural network can include acquiring and correcting position or proximity information. The corrected images and / or corrected position or proximity information are input into the neural network and processed to estimate the movement and positioning of the devices of the solar panel processing system, such as the movement and positioning of devices associated with autonomous vehicles, storage vehicles, robotic devices, and installation devices. The estimated movement and positioning are published to the control systems associated with the individual devices of the solar panel processing system or to the main controller of the entire solar panel processing system.

[0089] In some embodiments, the signal from the optical sensor can be input into a controller. In some embodiments, the solar panel processing system can further include an orientation assembly 804 (see Figure 8) configured to tilt the arm - end assembly tool 100 relative to the mounting structure 604. In such an embodiment, the controller 112 can control the orientation in response to an input from an optical signal that indicates that the solar panel being installed is not properly oriented or leveled relative to the mounting structure (e.g., torque tube 604). It will be appreciated that although the orientation assembly 804 is shown coupled to the force - torque sensor 110, those of ordinary skill in the art will readily recognize other means of implementing the orientation assembly 804.

[0090] In some embodiments, the controller 112 can also be configured to control the attachment device 104 to activate or deactivate its attachment / detachment. For embodiments where the attachment device 104 is a suction cup, a vacuum enables the coupling or release of the solar panel 120 to the arm - end assembly tool 100.

[0091] In some embodiments, the mounting structure 604 can have an octagonal cross - section, such as shown in Figure 6A 、 6B and FIGS. 7A - 7D, to form a torque tube that prevents the clamp assembly 602 from accidentally slipping. However, other cross - section shapes can also be used, such as square, oval, or other shapes. Additionally, the mounting structure 604 can use a circular cross - section shape.

[0092] In some embodiments, the assembly tool 100 can be configured to couple to an assembly mobile robot 903 (an example of which is shown in Figure 9 and Figure 10 . The assembly mobile robot 903 can be configured to position the arm - end assembly tool 100 relative to a stack or storage container 905 of solar panels, move a selected solar panel, and position the selected solar panel relative to the mounting structure 604. In some embodiments, the assembly mobile robot 903 can be operatively coupled to the arm - end assembly tool 100 via the force sensor 110 (or, where applicable, the orientation assembly 804). In some embodiments, the assembly mobile robot can also be operatively coupled to the controller, enabling an operator of the assembly mobile robot to control various functions of the arm - end assembly tool 100, such as activation and / or deactivation of the connection device 104, advancement and / or retraction of the clamping tool, and activation and / or deactivation of the engagement member relative to the clamp assembly.

[0093] Now refer to Figure 1 、 6A, 6B, 7A - 7D, 9, and 10. In operation, a solar panel 120 is obtained and positioned above the mounting structure 604. Then, the solar panel is tilted relative to the mounting structure 604 such that the leading edge of the solar panel (i.e., the edge that will be adjacent to the edge of a previously installed solar panel, or for the first solar panel, the edge that will be adjacent to a stop fixed to the mounting structure 604) is oriented closer to the mounting structure 604 than the opposite trailing edge. Then, the leading edge is placed in a receiving channel (a receiving channel positioned along the edge of a previously installed solar panel, i.e., as part of a clamp assembly, or a receiving channel in the stop), and the tilt of the solar panel is reduced to the mounting position on the mounting structure. As the solar panel is biased into the receiving channel, the tilt angle is reduced such that at this mounting position, the edge region of the top planar surface of the solar panel (i.e., the photovoltaic active surface facing the sun) is captured within the receiving channel. Figure 6A and 6B An exemplary embodiment of the receiving channel 610 on the clamp assembly 602 is shown in

[0094] Once the solar panel is in place on the mounting structure, the force - torque actuator 110 actuates the guide assembly 106 of the arm - end assembly tool 100 to bring the engaging member 108 - A of the clamping tool 108 into contact with the clamp assembly 602. The clamp assembly is initially positioned outside the area on the mounting structure to be occupied by the solar panel being installed, but is also close enough so that the relevant components of the arm - end assembly tool 100 can reach it. The surface and features of the engaging member 108 - A can be positioned and sized to match complementary features on the clamp assembly 602. After this contact, the force - torque actuator 110 is actuated (either continued actuation or actuation in a second mode) to axially slide the clamp assembly 602 along a portion of the length of the mounting structure 604. The axial sliding of the clamp assembly 602 causes the receiving channel of the clamp assembly 602 to engage with the trailing edge of the newly installed solar panel. Sensors, such as sensors in the force - torque actuator 110 or in the clamping tool 108, can provide feedback to a controller, thereby indicating the complete engagement of the receiving channel of the clamp assembly 602 with the trailing edge of the solar panel. Once the clamp assembly 602 is in place, the guide assembly 106 retracts, and the installation of the next solar panel can proceed.

[0095] In some embodiments, the linear guidance assembly 106 may include a proximity sensor 108-B configured to sense the distance between the engagement member 108 and the trailing edge of the solar panel 120 during the installation operation of the solar panel 120. The output from the proximity sensor 108-B may be used to appropriately control the speed of the gripping tool 108 during the operation of the linear guidance assembly 106 to avoid excessive force and impact on the solar panel 120. In some embodiments, the proximity sensor 108-B may be, for example, an optical or audio sensor (e.g., sonar) that detects the distance between the leading edge of the solar panel 120 and the engagement member 108; in other embodiments, the proximity sensor 108-B may be a limit switch that retracts upon contact.

[0096] Further reference Figure 9 and Figure 10 , the assembly mobile robot 903 can be implemented using a ground vehicle 907. For example, the ground vehicle 907 can be implemented as an electric vehicle (EV). The ground vehicle 907 can autonomously move adjacent to the mounting structure 604. Although not shown, the ground vehicle 907 can move along a track or guide rail attached to or separated from the mounting structure. In some embodiments, the ground vehicle 907 can be controlled using sensors, or based on input or feedback from sensors. For example, the sensors can be optical sensors or proximity sensors. In additional embodiments, a neural network using artificial intelligence can be used to control the movement of the ground vehicle 907, such as by analyzing the operating environment and formulating instructions for the movement of the ground vehicle.

[0097] Figure 10 An embodiment of a solar panel handling system having two robotic arms is illustrated, where two assembly tools are coupled to an assembly mobile robot using the respective robotic arms.

[0098] As Figure 9 shown, a storage container 905 containing a solar panel to be installed can be disposed on the ground vehicle. Here, Figure 9 An embodiment of a solar panel handling system including an arm assembly tool 100 is illustrated, where the arm assembly tool 100 is coupled to an assembly mobile robot using a robotic arm. Alternatively, as Figure 10 shown in Figure 10 An embodiment of a solar panel handling system having two robotic arms is illustrated, where two assembly tools are coupled to an assembly mobile robot using the respective robotic arms. In embodiments of the present disclosure, the robotic arm can be an articulated arm having two or more segments coupled using joints, or alternatively can be a truss arm. The illustrations herein are intended to disclose the use of any type of arm according to the present disclosure.

[0099] According to Figure 9 , for example, the robotic arm of the arm assembly tool 100 having an upper section 908 and a lower section 909 can provide increased operational flexibility while maintaining light weight and simple operation. As Figure 9 additionally shown in, the second robotic arm 911 can be provided with the arm assembly tool 100 having a nut tightener or nut driver at its end to fix the solar panel to the mounting structure 604. Although any type of robotic arm can be used for the second robotic arm 911, Figure 9 an example of using an articulated arm having a nut tightener or nut driver at its end is illustrated. Here, the robotic arms 100 and 911 can autonomously operate using computer vision with neural networks and artificial intelligence control. Alternatively, the robotic arms 100 and 911 can be manually operated or remotely controlled.

[0100] In some embodiments, the ground vehicle 907 can be an autonomous vehicle where neural networks and artificial intelligence control the movement and operation, and the modular vehicle 1005 is towed or coupled to the ground vehicle 907. In other embodiments, the modular vehicle 1005 can be an autonomous vehicle where neural networks and artificial intelligence control the movement and operation, and the ground vehicle 907 is towed or coupled to the modular vehicle 1005. Additionally, in some embodiments, the assembly mobile robot 903 is also mounted on one of the ground vehicle 907 and the modular vehicle 1005. In other embodiments, the assembly mobile robot 903 can be mounted on a dedicated robotic vehicle.

[0101] Figures 11A to 11C A process for installing solar panels is shown in. As Figure 11A shown, a pallet of solar panels can be transported by a truck. In some embodiments, the pallet can form the storage container 905 for the solar panels. The pallet can include machine-readable markings, such as barcodes, QR codes, or other manufacturing references, which can be read to provide information about the solar panels, installation instructions, or other information to be used during the installation process, especially information to be used by neural networks and artificial intelligence control. For example, this information can include the number of solar panels, the type of solar panels, the physical characteristics of the solar panels (e.g., dimensions), the characteristics related to installation (e.g., hardware type and location), installation instructions, or other characteristics of the solar panels, the storage of the solar panels on the pallet, and information related to installation. Additionally, using the machine-readable markings, the system can control the supply or replenishment of the panel boxes in the correct order and / or ensure the use of panels with similar impedance from the factory.

[0102] As Figure 11BAs shown, a mechanized device such as a forklift can be used to move and position the pallet on the ground vehicle. Here, the forklift can be manually operated, remotely operated, or self-driving. In Figure 11B the pallet is located on the ground vehicle. Alternatively, the pallet can be located on the modular vehicle. Then, as Figure 11C shown, the arm of a robot is used to install the solar panel. In the illustrated example, two arms are used to handle the respective solar panels to be installed on the respective mounting structures. Here, the ground vehicle moves between the two respective mounting structures. In addition, a modular vehicle is provided that can be separated from the ground vehicle.

[0103] As those of ordinary skill in the art will recognize, modifications and variations in the embodiments can be used. For example, as Figure 12A and 12B shown, two modular vehicles can be provided for the respective robotic arms. In another alternative, the modular vehicle can be connected to the ground vehicle instead of being separated. Thus, as Figure 12A shown, the robotic arm can engage the respective solar panel to be installed, as Figure 12B shown.

[0104] In some embodiments, as Figure 13 shown, computer vision registration can be used to achieve the installation. For example, as mentioned above, optical sensors and the like can be used together with a neural network for artificial intelligence.

[0105] In some embodiments, as Figure 14 shown, if the modular vehicle is used with the ground vehicle, when all the solar panels of the modular vehicle are installed, the modular vehicle can be exchanged with a supplementary modular vehicle. Here, the computer vision process can be used to communicate with and control an autonomous and independent vehicle (e.g., a forklift) to bring an additional solar panel box. Thus, the supply of solar panels can be replenished.

[0106] In the supplementary operation of the forklift example, the forklift (whether autonomous, remotely controlled, or manually operated) can be used to return the empty box or container of the solar panel to the waste area, remove the strap, open the lid, or cut off the face of the box being transported, pick up the box to correct the rotation / orientation of the solar panel, or other tasks. In addition, the forklift can stay near the ground vehicle to wait for the system to deplete the next box of solar panels. Thus, the forklift can manually or automatically discard the depleted box, position the next box on the ground vehicle or the modular vehicle, open the box (including removing the strap, opening the lid, or cutting off the face of the box), and move back away from the ground vehicle / modular vehicle. As described, for example, this replenishment can be autonomous, remotely controlled, or manually operated.

[0107] Figures 15 - 34Provides a detailed illustration of an example configuration of a system for installing solar panels according to an embodiment of the present disclosure.

[0108] Computer vision and artificial intelligence (AI) technologies

[0109] According to some embodiments, computer vision and AI technologies can be used to determine positions along an installation structure (e.g., a torque tube) where solar panels can be installed. The installation positions of the panels can be based on the positions of previously installed panels. Some embodiments use the six degrees of freedom (6DoF) pose of previously installed panels. The term "6DoF" represents six degrees of freedom, namely three rotational axes (yaw, pitch, and roll) and three translational axes (x, y, and z). The estimation of the 6DoF pose of a panel includes the positions (x, y, z) in 3D space of points (or key points) on the panel, such as the corners of the frame or some other visual reference, as well as the rotation angles (rx, ry, rz) of the panel about the three Cartesian axes. In some embodiments, computer vision or AI is used to estimate the 6DoF pose of a previously installed panel to determine the position along the torque tube where the next panel is to be installed. The next panel to be installed along the torque tube typically has a certain constant offset from the positions (x, y, z) of the key points of the previously installed panel, while having the same rotation angles (rx, ry, rz).

[0110] Estimation of 6DoF poses for panel placement and panel picking

[0111] Although some of the embodiments described herein relate to panel placement (e.g., mounting along a torque tube), those embodiments rely on techniques that are also applicable to panel picking, e.g., picking a panel from a storage location such as a module bin or cradle. For panel picking, as with panel placement, the 6DoF pose of the panel in question is determined. For panel placement, the panel in question may be a previously mounted panel, as described above. In contrast, for panel picking, the panel in question may be the current (or outermost) panel (in the storage location) to be picked for mounting. For panel placement, the 6DoF pose of the previously mounted panel may be used to determine the (offset) position of the next panel to be mounted along the torque tube, as described above. In contrast, for panel picking, the 6DoF pose of the current panel to be mounted may be used to determine the pick point (typically the center point) of the panel to be picked. For panel picking, one challenge is to pick a panel such that the EOAT is centered with respect to the panel. This centering of the EOAT with respect to the picked panel ensures an even load distribution on the EOAT. For panel placement, as described above, the challenge is to mount the next panel such that it has a specific pose (typically: x, y plus the panel width plus [e], z, rx, ry, rz) with respect to the previously mounted panel. However, for both panel placement and panel picking, the goal of CV / AI is to determine the 6DoF pose of the panel in question. And this 6DoF pose can be determined through techniques and technologies based on the techniques described herein.

[0112] In the following description, the term "computer vision" is used to refer to "classical" computer vision techniques and algorithms, and the term "artificial intelligence (AI)" is used in place of "machine learning (ML)" or "deep learning (DL)" because the present disclosure generalizes well to current and newly developed technologies.

[0113] The term "image" is used to include not only camera images, but also data and images from other types of sensors, such as time-of-flight sensors, LiDAR, or other sensors that image or scan a field of view (FoV); and the terms "preprocessing" and "postprocessing" may involve different hardware (e.g., computing devices, sensors) for what is initially or subsequently processed.

[0114] Exemplary non-AI methods for pose estimation

[0115] In some embodiments, non-AI techniques are used to estimate the 6DoF pose. Different sensors (e.g., stereo cameras, LiDAR) and computer vision algorithms may be used to determine the 6DoF pose of the panel. These different methods are described below as different preprocessing and postprocessing methods for AI processing. These preprocessing and postprocessing steps themselves (without AI) may be combined to form a computer vision pipeline that can be used to determine the 6DoF pose of the panel.

[0116] Exemplary AI Models and Inferences

[0117] Now turning to different types of AI models and inferences applicable to determining the 6DoF pose of a solar panel. AI-based inferences may be based on bounding boxes, segmentation, key points, depth, and / or 6DoF pose. Each of these different techniques is discussed in turn below.

[0118] Segmentation

[0119] Segmentation in AI refers to the process of dividing an image or video into meaningful and semantically coherent regions. The goal of segmentation is to partition visual data into distinct regions based on common characteristics of the visual data, such as color, texture, or object boundaries. In computer vision, traditional segmentation techniques include methods such as thresholding, region growing, edge-based segmentation, and clustering algorithms such as k-means or mean shift. These methods rely on manual features and heuristics to segment images.

[0120] In AI, deep learning methods, particularly convolutional neural networks (CNNs), have shown exceptional performance in segmentation tasks. Fully convolutional networks (FCNs), U-Net, Mask R-CNN, and DeepLab are popular architectures for semantic and instance segmentation. These models leverage their ability to learn and extract complex features from images, enabling accurate and efficient segmentation.

[0121] Different types of segmentation techniques can be used in AI, including semantic segmentation and instance segmentation. Semantic segmentation involves labeling each pixel in an image or video frame with a corresponding class label. The output is a per-pixel classification map where each pixel is assigned a semantic category or class label. Semantic segmentation focuses on capturing the semantic meaning of a scene and is used for scene parsing, object recognition, and high-level understanding. Instance segmentation goes beyond semantic segmentation and aims to separate and identify individual objects within an image. It assigns a unique label or identifier to each pixel belonging to a specific object instance. In instance segmentation, each object is segmented individually, allowing for precise delineation and separation of object boundaries. This technique is used for object detection, tracking, counting, and detailed object analysis. For the purposes of the discussion below, the term "segmentation" refers to instance segmentation (unless otherwise specified).

[0122] In some embodiments, segmentation is used to determine the 6DoF pose of the panel. Taking panel segmentation as input, various computer vision algorithms can find the pixel positions of the four corners of the panel frame in the image, and then these positions are used as input to a Perspective-n-Point solver to determine the 6DoF of the panel in 3D space (or "world space").

[0123] In some embodiments, based on the segmentation, the pixel positions of the panel corners are determined by one or more of various computer vision techniques, such as the Hough transform, the Ramer–Douglas–Peucker algorithm, or some other computer vision post-processing. For example, the Hough transform can produce four Hough lines that fit to the four edges of the segmentation, which in turn correspond to the four edges of the panel. And the four intersection points of these four Hough lines produce the four corners of the panel. Figure 51 An exemplary application 5100 of the Hough transform for identifying the corners of a solar panel 5110 according to some embodiments is shown. The intersection of Hough lines 5102 and 5106 produces corner A, the intersection of Hough lines 5102 and 5108 produces corner D, the intersection of Hough lines 5104 and 5106 produces corner B, and the intersection of Hough lines 5104 and 5108 produces corner C.

[0124] In some embodiments, the Ramer-Douglas-Peucker algorithm is used to determine the pixel positions of the four panel corners. In some embodiments, the algorithm is restricted to produce a four-sided polygon approximation. For example, in the OpenCV implementation of the algorithm (the ApproxPolyDP() function), the parameter epsilon can be optimized to produce a four-sided polygon approximation, for example, based on a binary search of epsilon. As will be appreciated by those skilled in the art, other suitable computer vision algorithms can also be used for the post-processing of the segmentation to find panel corners or other key points. In some embodiments, these key points (e.g., panel corners) are used as inputs to a Perspective-n-Point solver to determine the 6DoF pose of the panel.

[0125] In some cases, the above post-processing methods of fitting a single line along the length direction of the panel edge (e.g., the Hough transform, the Ramer-Douglas-Peucker algorithm) may produce inaccurate corner positions, whether at the far half or the near half of the panel. The reason is that the solar panel is not a perfectly rigid body. For example, when mounted on a torque tube, the panel overhangs the tube and deflects or deforms under its own weight. Therefore, any approximation of the longitudinal edge by a single line may be inaccurate for one corner or the other. To compensate for this inaccuracy, in some embodiments, instead of fitting a single line, two lines are fitted to the longitudinal edge of the panel, one line for the far half of the panel and one line for the near half of the panel. Then, the Perspective-n-Point solver can solve for the 6DoF pose (e.g., where z is the height change due to panel deformation) based on the coordinates (x, y, z) of the panel corners that deviate from a perfectly rigid body.

[0126] Another post - processing method that may be particularly suitable for AI - based segmentation is a method of analyzing the pixel intensity of the shadow between adjacent panels. AI - based segmentation may perform worse on images with multiple panels (compared to images with a single panel). For example, for an image with multiple panels, the segmentation of the panel in question may retreat into the interior of the panel (rather than coinciding with the outer edge of the panel frame as expected).

[0127] Such a difference can be observed along the length of the panel in question near the adjacent panel. To compensate for this difference, in some embodiments, the shadow between the panel in question and the adjacent panel is used as a visual reference from which the true edge of the panel in question is identified. Across different lighting conditions, the shadow between adjacent panels is largely consistent and distinct. In some embodiments, computer vision algorithms such as binarization, thresholding, and Hough transform are used to identify the true edge of the panel in question.

[0128] Depth Estimation

[0129] In some embodiments, depth estimation is used to determine the 6DoF pose of a panel. Depth estimation is the determination of the distance of a given object within the field of view (FoV) relative to a sensor (e.g., a camera, LiDAR, time - of - flight sensor).

[0130] Depth estimation can be understood as a kind of segmentation. The object in question is segmented or distinguished from the (more) distant background. Thus, the panel appears as a connected segmentation in a disparity map (using computer vision) or depth visualization (using AI). Figure 52 An exemplary application 5200 of segmentation according to some embodiments is shown. The result of the segmentation can be used as an input for post - processing to find panel corners or other key points for the purpose of 6DoF pose estimation, as described above.

[0131] Using AI, for example, a neural network for monocular depth estimation (MDE) can be used to perform depth estimation. The following reference is a comprehensive survey of the state - of - the - art methods for MDE, which is incorporated herein by reference: “The monocular depth estimation challenge” by J. Spencer et al., Proceedings of the IEEE / CVF Winter Conference on Applications of Computer Vision (pp. 623 - 632) (2023).

[0132] Using computer vision, depth estimation can be performed using two cameras (or a stereo camera) to generate a disparity map based on corresponding points and the baseline distance between the two cameras. The disparity or perspective difference between the images from the two (stereo) cameras indicates the distance of the object from the cameras. Using CV, depth estimation can have problems under certain lighting or surface conditions (e.g., textureless regions and specular reflections on panels). Conventional techniques such as photometric consistency methods and active stereo vision (using IR structured light projection) can be employed to improve the accuracy of depth estimation. For example, multi-baseline stereo vision involving three or more cameras (with multiple baseline distances between the cameras), such as a three-eye or four-eye configuration, and associated algorithms (e.g., semi-global matching and iterative matching algorithms) can also be used to improve the resolution and accuracy of depth estimation. The following references are a comprehensive survey of state-of-the-art methods for multi-baseline stereo vision, which are incorporated herein by reference: H. Hirschmuller, “Stereo Processing by Semiglobal Matching and Mutual Information”, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 30, no. 2, pp. 328 - 341, February 2008, doi: 10.1109 / TPAMI.2007.1166; S. Patil et al., “A Comparative Evaluation of SGM Variants (including a New Variant, tMGM) for Dense Stereo Matching”, ArXiv, abs / 1911.09800 (2019).

[0133] Whether using AI or CV, depth estimation (as a type of segmentation) uses the same post-processing (as described above) for the purpose of locating the panel corners in the image. For example, the determination of the corner positions (whether by depth estimation or segmentation) can be corrected or refined (before being used as the input to a PnP solver) by the so-called “two-pass homography” post-processing described below.

[0134] Key point detection

[0135] Some embodiments use a type of AI-based inference called keypoint detection to determine the 6DoF pose of the panel. Keypoint detection is a fundamental task in computer vision and AI, which involves identifying and locating specific points or landmarks in an image or video. These keypoints represent unique features in the visual data, such as corners, edges, or other regions of interest. The goal of keypoint detection is to accurately locate and describe these keypoints, enabling various applications such as object recognition, tracking, pose estimation, and image alignment.

[0136] Keypoint detection typically includes the following steps or algorithms. Preprocessing: The input image or video frame is usually preprocessed to enhance its quality and reduce noise. Common preprocessing steps include resizing, normalization, and grayscale conversion. Multiple algorithms can be used for keypoint detection, depending on the specific requirements and characteristics of the data. Some embodiments use corner detection: algorithms such as the Harris corner detector or the Shi-Tomasi corner detector to identify corners in the image based on local intensity changes. Some embodiments use scale-space extremum detection, including methods such as Difference of Gaussians (DoG) or Laplacian of Gaussian (LoG), to detect keypoints at different scales by finding local extrema in the scale-space representation of the image. Some embodiments use interest point detectors, which include algorithms such as SIFT (Scale-Invariant Feature Transform) or SURF (Speeded-Up Robust Features), to identify keypoints based on local image gradients and their responses to different scales and orientations. Some embodiments use deep learning-based methods, such as Convolutional Neural Networks (CNNs), which can be trained to directly detect keypoints by learning from annotated datasets.

[0137] Models such as DenseNet, CornerNet, or OpenPose use deep learning techniques for keypoint detection. Once the keypoints are detected, their precise positions within the image need to be determined. This process may include refining the initial detection, which is achieved by using techniques such as sub-pixel interpolation or optimization algorithms to improve the localization accuracy.

[0138] The keypoints can be the four corners of the solar panel frame, as described above. The keypoints can also be any visual reference provided by a single panel, such as a common grid pattern, as discussed below. The keypoints can also be points on the solar panel, such as the center point.

[0139] Regarding the segmentation model / inference, the keypoint detection model / inference can use their own post - processing computer vision methods. For example, the Composite Of Shifted Filter Responses (COSFIRE) filter can be optimized for pattern recognition of corners to distinguish them from the background. The keypoint detection model may generate a rough Region of Interest (ROI) for input into the COSFIRE filter, which in turn produces a more fine - grained corner detection.

[0140] As recognized by those skilled in the art, other applicable computer vision algorithms and techniques can also be used for post - processing of keypoint detection to find panel corners for the purpose of 6DoF pose estimation.

[0141] Bounding box

[0142] Some embodiments use bounding boxes (another AI - based inference technique) to determine the 6DoF pose of a panel. Bounding box inference in AI refers to the process of predicting and localizing an object by drawing a rectangular bounding box around the object in an image or video. The bounding box provides an approximation of the position and extent of the object within the visual data. This technique is widely used in object detection, localization, and tracking tasks. Bounding box inference typically relies on object detection models that have been trained using machine learning algorithms. Popular object detection models include Single Shot MultiBox Detector (SSD), You Only Look Once (YOLO), and Faster R - CNN (Region - based Convolutional Neural Network). These models are designed to detect and classify objects in an image or video frame. The object detection models are trained on a labeled dataset, where each object of interest is annotated with a bounding box that tightly encloses the object. The training data also includes the class label corresponding to each object category. During training, the model learns to identify objects and predict their bounding boxes based on visual features extracted from the input data. During the inference phase, the trained object detection model takes an input image or video frame and processes it to detect objects and predict their bounding boxes. The model analyzes the visual features and makes predictions about the presence, category, and location of objects within the image. The object detection model predicts the coordinates of the bounding box for each detected object. The bounding box is typically represented by four values: the x - coordinate and y - coordinate of the top - left corner, and the width and height of the box. These values are used to draw a rectangle around the object, indicating its estimated position in the image or frame. Since multiple bounding box predictions can overlap or enclose the same object, a post - processing step called non - maximum suppression is typically applied. Non - maximum suppression aims to remove redundant or duplicate bounding boxes, thus ensuring that only the most accurate and reliable bounding box is retained for each object. This step helps to eliminate duplicate detections and improve the accuracy of bounding box inference.

[0143] In some embodiments, a bounding box is used to identify a given solar panel by bounding or encompassing the entirety of the solar panel within an image. The bounding box can be understood as a special case of keypoint detection for the purpose of determining the 6DoF pose of the panel. The reason is that given a typical top-down perspective, two corners of the bounding box may coincide with two physical corners of the panel frame. Specifically, the upper right corner and the lower left corner of the bounding box coincide with the far right corner and the near left corner of the panel frame. This correspondence allows the bounding box reasoning to be interpreted as keypoint detection (of two corners). This correspondence allows the bounding box model to replace the keypoint detection model in methods that do not require identifying all four panel corners, and such a model replacement may be beneficial for training and / or inference.

[0144] Whether as keypoint detection or bounding box, the AI-based reasoning for determining keypoints (e.g., panel corners) can use the same post-processing (e.g., the "two-pass homography" algorithm described below) for the purpose of determining the 6DoF pose of the panel.

[0145] Six degrees of freedom

[0146] Six degrees of freedom (6DoF) pose estimation refers to the task of using a machine learning model to estimate the position and orientation of an object in 3D space. The term "6DoF" represents six degrees of freedom, namely three rotational axes (yaw, pitch, and roll) and three translational axes (x, y, and z). Pose estimation is crucial in various computer vision applications such as robotics, augmented reality, and object tracking, where accurate knowledge of the position and orientation of an object is essential.

[0147] Some embodiments use an AI model for 6DoF pose estimation. The AI models can be classified into two main types: feature-based methods and direct regression methods.

[0148] First, feature-based methods involve extracting keypoints or features from an object or scene and matching them between a 3D model and an input image. Then, the pose is estimated based on the spatial relationships between the matched 3D-2D feature correspondences. Traditional computer vision techniques or deep learning-based methods can be used for feature extraction and matching. Traditional feature-based methods use techniques such as SIFT (Scale-Invariant Feature Transform), ORB (Oriented FAST and Rotated BRIEF), or SURF (Speeded-Up Robust Features) to detect and match keypoints between a 3D model and a 2D image.

[0149] Then, the RANSAC (Random Sample Consensus) or PnP (Perspective-n-Point) algorithm is used to estimate the 6DoF pose based on the matched correspondences. In deep learning-based feature-based methods, instead of using handcrafted features, deep learning models such as CNN can be used to directly learn the feature representation from the input image. PoseNet and PoseCNN are examples of deep learning-based feature-based methods that use CNN to predict the 6DoF pose.

[0150] Second, direct regression methods directly predict the 6DoF pose parameters (translation and rotation) from the input image, thus eliminating the need for feature extraction and matching. CNN-based regression models use the CNN architecture to directly regress the 6DoF pose parameters from the input image. The network takes the image as input and outputs the pose values as continuous numbers. Models such as DeepIM and PVNet are examples of CNN-based regression methods. Some hybrid methods combine feature-based and direct regression techniques. They use CNN to predict an initial pose estimate and then use a feature-based method to refine the pose for higher accuracy.

[0151] If an AI model for 6DoF inference is accurate enough, it can be used to determine the 6DoF pose of the solar panel (without post-correction processing). However, if the 6DoF model is not accurate enough, these models can still provide information that can be interpreted according to other types of inferences, such as segmentation, keypoint detection, and bounding boxes, which can then be corrected or refined through post-processing. For example, an initial (coarse) 6DoF estimate of the panel pose can be used as a coarse identification of the corners for the purpose of keypoint detection or bounding box inference. For example, the initial 6DoF pose estimate can also be used to (coarsely) identify the panel or segment the panel from its background.

[0152] In summary, different types of AI models and inferences applicable to determining the 6DoF pose of solar panels have been described above. These AI models can be zero-shot, one-shot, or few-shot and can be trained differently, for example, using real images or synthetic images, or using supervised learning or unsupervised learning.

[0153] In some embodiments, each AI model or inference uses its own post-processing method (e.g., analyzing the shadows, filters between adjacent panels). In some embodiments, different AI models / inferences use the same post-processing method (e.g., the "two-pass homography" described below).

[0154] AI models can be combined such that the output of one model serves as the input to another model. For example, a first model (e.g., YOLO) outputs bounding boxes, which in turn serve as the input to a second model (e.g., YOLO, SSD, RCNN, SAM) for segmentation. A single type of model can be used twice, where a second instance of the same model takes the output of the first instance as input. For example, YOLO for bounding boxes can be used as the input to YOLO for keypoint detection. An AI pipeline can include any number of AI models or integrated AI models. Ensemble learning in machine learning refers to the technique of combining multiple individual models (referred to as base models or weak learners) to form a more powerful and accurate model (referred to as an ensemble model). The main idea behind ensemble learning is that by aggregating the predictions of multiple models, the ensemble can make more reliable predictions than any single model alone. There are several ensemble methods for combining the predictions of base models. A base model is each individual model that forms the ensemble. They can be any machine learning algorithm, such as decision trees, random forests, support vector machines, neural networks, or any other model. Each base model is trained on a subset of the training data, or some variation is introduced to create diversity among the models. For example, in a voting-based ensemble, each base model makes predictions independently, and the final prediction is determined based on a majority vote (for classification problems) or an average (for regression problems) among the predictions. Bagging (Bootstrap Aggregating) involves training multiple base models on different random subsets of the training data with replacement. The final prediction is obtained by averaging (for regression) or voting (for classification) the predictions of the individual models. Boosting algorithms, such as AdaBoost, Gradient Boosting, or XGBoost, train base models sequentially, where each subsequent model focuses on the instances that the previous model had difficulty with. The predictions of all the models are combined to form the final prediction, usually through weighted voting. Stacking combines the predictions of multiple base models by training a meta-model that learns to make predictions based on the outputs of the individual models. The predictions of the base models are used as features, and the meta-model is trained on this augmented dataset.

[0155] The strength of an ensemble lies in the diversity among its base models. Diversity is achieved through various means such as using different algorithms, altering the model architecture, training on different subsets of data, or introducing randomness during training. Diverse models make different types of mistakes, and when combined, they can compensate for each other's weaknesses, leading to improved overall performance. Ensemble models typically outperform individual models because they can capture different aspects of the data and combine their strengths to make more accurate predictions. Compared to individual models, ensemble models are generally more robust to noise and overfitting because the mistakes made by some models can be compensated by other models. By reducing the impact of bias and error of individual models, ensemble models can generalize to unseen data, resulting in better performance.

[0156] The similarities or overlapping applicability between different types of AI models / inferences are described above, allowing one type of model / inference to be interpreted as a special case of another (more general) type of model / inference. Such generality results from the special characteristics of the imaging scenarios in the use cases of the present invention. Conversely, these different types of models / inferences admit the same post-processing (e.g., the "two-pass homography" algorithm) due to their generality.

[0157] The ways in which different AI models / inferences can be combined with different computer vision algorithms are described below. Computer vision algorithms can be integrated within the AI pipeline as post-processing or pre-processing. Computer vision algorithms can also be integrated within the AI pipeline based on an architecture that allows for a deeper complementarity between computer vision and AI, thus fully leveraging their respective strengths. Computer vision algorithms can also be combined on their own without using AI.

[0158] Incidentally, the deeper complementarity between the disclosed CV and AI exemplifies what O’Mahony et al. (2020) describe as “mixing hand-crafted approaches with [deep learning] DL for better performance”. As they observe, “there is a clear trade-off between traditional CV and deep learning-based methods. Classical CV algorithms are mature, transparent, and optimized for performance and energy efficiency, while DL offers higher accuracy and generality at the cost of significant computational resources.” However, as O’Mahony et al. observe, in applications where DL is (still) not yet perfect (e.g., 3D vision), CV and AI / DL can be fruitfully combined. See N. O’Mahony et al., “Deep learning vs. traditional computer vision”, Advances in Computer Vision: Proceedings of the 2019 Computer Vision Conference (CVC), Springer Nature Switzerland AG, Volume 11 (pp. 128-144) (2020).

[0159] Whether as post-processing or pre-processing, whether as part of a deeper integration with AI or as a stand-alone computer vision pipeline without AI, computer vision algorithms rely on invariant structures that are distinguishable across different images. Such invariant structures can be attributed to physical structures, such as the grid pattern formed by photovoltaic cells and their electrical connections. Such invariant structures may also be attributed to the immaterial structure of illumination, such as the shadows cast between adjacent solar panels as a kind of negative space. Whether physical or immaterial, whether as positive or negative space, the invariant structures that support CV algorithms provide a consistent visual pattern against which the visual pattern can be analyzed for its semantic content.

[0160] Post-processing using the “two-pass homography” algorithm

[0161] Post-processing that compensates for or corrects inaccuracies in (AI-based or computer vision-based) segmentation is called the “two-pass homography” (TPH) algorithm. This algorithm relies on a regular visual pattern or fiducial on the solar panel, such as a grid pattern. This grid pattern (or other visual fiducial) can be used to compensate for / correct inaccuracies in segmentation because it provides reference points by which the pixel positions of the panel corners can be indirectly inferred through a homography transformation. In some embodiments, the TPH algorithm includes the following steps:

[0162] (1) Use panel segmentation as a mask to filter out the background in the image (leaving only the foreground panel).

[0163] (2) Use the Shi-Tomasi algorithm to identify the grid intersections in the masked panel from (1) as "corners".

[0164] (3) Use the "corners" from (2) to calculate the homography matrix H 1 .

[0165] (4) Based on H from (3) 1 , determine the positions (described below) that the grid intersections from (2) should be in the millimeter space.

[0166] (5) Determine the homography matrix H 2 to transform the grid intersections from (2) in the image space to their positions in the millimeter space in (4).

[0167] (6) Use the inverse of H from (5) 2 to back-project the corners in the millimeter space to the image space.

[0168] Note that the term "corner point" is a term in the field of computer vision and generally refers to a point of interest where the intensity gradient is maximum in all directions. "Corner points" (or italicized corner points) should be distinguished from physical corners (e.g., the corners of a panel frame).

[0169] For (1), the filtering operation may be as simple as a binary AND between the mask and the original image.

[0170] For (2), other corner point finding algorithms (e.g., Harris-Stephens algorithm) can be used. Note that the grid intersections or "corner points" found in (2) may not exactly coincide with the intersections of the grid lines. The reason is that the grid intersections may include various shapes (e.g., diamonds, two triangles with a common vertex), and their gradient intensities can cause the corner point detection algorithm to detect "corner points" that are off-center, an example (5300) of which is shown in Figure 53 . However, it can be said that as the number of detected "corner points" increases as input for homography calculation, the (random) eccentricity of the detected "corner points" is canceled out. In Figure 53 , more distinct and thicker grid lines (sometimes called "main" grid lines) are useful for the TPH algorithm. Various computer vision algorithms (e.g., Gaussian blur with a kernel size of 5 for example) can be used to distinguish the "main" grid lines from the (darker) "secondary" grid lines.

[0171] In some embodiments, the AI ​​model can also be used to find grid intersections. The AI ​​model can use the original image or a homography matrix H based on the (“first pass”) 1 The projected (or distorted) image is taken as input. 1 In space, the panel is projected to an approximate top view. From this view, the AI ​​model may be easier to train because perspective changes are minimized. However, if the AI ​​model takes the original image as input, and if inference produces all grid intersections, then the "two-pass homography" algorithm simplifies to a "single-pass" algorithm. The reason is that for all grid intersections detected by the AI ​​model, as described below, the association in (4) simplifies to a simple enumeration (i.e., the relative positions of the grid intersections are the same).

[0172] For (3), the (“first pass”) homography matrix H 1 Panel projections that are not perfectly square in "millimeter space" may be produced. The term "millimeter space" refers to a 2D projection space where the four (physical) corners of a solar panel have the following (x,y) coordinates: (0,0), (0,L), (W,0), (W,L), where W and L are the width and length of the solar panel, respectively. However, the homography matrix H 1 should be accurate enough to allow associating the grid intersections found in (2) (in H 1 space) to where they should be in millimeter space.

[0173] For (4), the association (in H 1 The distance between the grid intersections found in (2) and where they should be in millimeter space) may be based on Euclidean distance (e.g., fraction of the shortest cell dimension such as width) or some other measure. For example, a Euclidean distance of a fraction of the shortest cell dimension (such as cell width) can ensure that the associations (between the grid intersections found in (2) and where they should be in millimeter space) are unique. An association Euclidean distance that is too large (e.g., greater than about half the width or height of each cell) may result in ambiguous associations that are not unique or involve associations of adjacent panels.

[0174] For (5), the (“second pass”) homography matrix H 2 Recalculate based on the more accurate (associated) grid intersection positions in millimeter space determined in (4).

[0175] For (6), the more accurate homography matrix H 2 It is used to back-project the corners with coordinates (0,0), (0,L), (W,0), (W,L) in millimeter space to image space. Figure 53 Indicated by crosshairs.

[0176] Figure 54Schematic diagram of detecting 5400 grid intersections 5402 using the two - pass homography algorithm described herein according to some embodiments.

[0177] Some embodiments use a TPH algorithm in the following form, which is generalized beyond grid intersections for using any panel feature as a visual reference as follows:

[0178] 1. Estimate the four corners k of the panel in the image space i .

[0179] 2. Find the homography matrix H from k i to the millimeter space, i.e., (0,0), (H,0), (H,W), (0,W). 1

[0180] 3. Find p i and q i , where p i is the pixel position of the panel feature (as a reference) in the image space, and q i is its corresponding position in the millimeter space based on H 1 .

[0181] 4. Find the homography matrix H from p i to q i . 2

[0182] 5. Based on H 2 -1 back - project the corners (0,0), (H,0), (H,W), (0,W) from the millimeter space to the image space.

[0183] Note that in the above general formula, p i and q i can be not only grid intersections, but also, for example, the centroid of a rhombus or the common vertex of two triangles. And p i and q i may be some other pixel structure, not just "corner points".

[0184] Also note that in the above general formula, the association of the corresponding source and target pixel positions as given in (3) and (4) ultimately still based on the four corners (not "corner points") k of the panel i (as given in (1) and (2)) as the basic reference. Depending on the geometry and features of the solar panel, other references can be used.

[0185] Pre - processing using multi - baseline stereo vision

[0186] As described above, depth estimation can be used for segmentation. Depth estimation can also be used to determine a bounding box that encompasses the entire extent of the panel. In this way, the bounding box can be used to extract or segment a region of interest (ROI) from the original image for input into an AI pipeline.

[0187] Figure 55 A schematic diagram of ROI segmentation 5500 according to some embodiments is shown. For example, ROI segmentation is performed on the original (full-size) images from two 12MP RGB cameras (5502 and 5504) before passing the (segmented) image as input into the AI pipeline 5506.

[0188] As a preprocessing step, ROI segmentation can be based on any of the following:

[0189] ·d 1 : A disparity map from single-baseline stereo vision.

[0190] ·d 2 : A disparity map from multi-baseline stereo vision.

[0191] ·d 3 、d 4 : Monocular depth estimation from a neural network.

[0192] ·f: Any (weighted) combination of the above.

[0193] ·m: Onboard processing that runs a bounding-box-as-ROI inference (e.g., the processing provided by Luxonis OAK-D and Stereolabs ZED cameras).

[0194] Figure 55 It is not a diagram of the design, but a diagram of the design space. Figure 55 Captures the preprocessing design space described above, which depicts the maximized design of the elements. Not all elements need to be used in practice. This design space allows for the integration (h) of structures in motion (accompanying the movement of the EOAT), AI models (Model 1,..., Model n), and a distributed robot operating system (ROS) architecture (e.g., ROS Node i,... ROS Node j).

[0195] Sensor Fusion with LiDAR and Monocular AI

[0196] Embodiments of preprocessing and postprocessing monocular AI (i.e., using images from a single camera) using computer vision techniques were described above. As described below, in some embodiments, computer vision is used not only for the preparation or correction steps, but also in the determination of 6DoF pose. Computer vision can be more tightly integrated (or built-in) with AI.

[0197] Figure 56 Illustrates the multi-data channel capabilities 5600 of LiDAR used in some embodiments. Range 5602 is the distance of a point from the lidar camera, which is calculated by using the time of flight of a laser pulse. Signal 5604 is the intensity of the laser returned from an object (usually represented by the color of the point cloud). Environment 5606 is the camera return that captures the intensity of ambient light at a predetermined wavelength (e.g., 865 nm wavelength). Reflectivity 5608 is the reflectivity of the surface (or object) detected by the LiDAR sensor. Some embodiments use LiDAR that provides access to data or images of reflectivity, ambient near-infrared (NIR), and (2D projection) range, as well as LUT mapping between these images. The same pixel sensor can be used for various data channels. A single LUT or lookup table (which captures the physical dimensions of the hardware) may be sufficient. External calibration may not be required, and there is no need to inquire whether the calibrated mapping is universal.

[0198] The LUT mapping allows for smooth analysis to transfer from one image space to another, such as from an NIR image to a point cloud.

[0199] Figure 57 Is a schematic diagram of an exemplary computer vision / AI architecture 5700 for determining the 6DoF pose of a solar panel according to some embodiments. The CV / AI pipeline 5700 takes two types of data or images as input from LiDAR: an ambient NIR grayscale image 5702 and range data 5706 in the form of a point cloud. Transformation 5704 transforms the NIR into range data. First, the ambient NIR image 5702 is used as input to the CV / AI pipeline (preprocessing 5708 and AI 5710), which locates the coordinates (x pixel , y pixel ) 5718 of key points (e.g., corners) on the panel in the image space. The CV / AI pipeline can include any of the preprocessing, AI models (e.g., segmentation, key point detection, bounding box), and postprocessing described above. Additionally, the point cloud range data 5706 is used as input to the CV pipeline, which identifies the panel by plane segmentation 5712 (i.e., fitting a plane to the points in the point cloud). For example, the fitted plane obtained in Hessian normal form provides the rotation angles (rx, ry, rz). Additionally, the coordinates (x pixel , y pixel ) 5718 of the relevant key points in the NIR image space are transformed (T) or mapped (through the LUT, as described above) (5714) to the point cloud to produce the coordinates (x, y, z) of the key points in world space. The combined output (x, y, z, rx, ry, rz) is the 6DoF pose of the panel in question.

[0200] Figure 58 is a schematic diagram of another exemplary computer vision / AI pipeline 5800 for determining the 6DoF pose of a solar panel according to some embodiments. This architecture is similar to the one described above with reference to Figure 57 except that pipeline 5700 relies only on data or images from the LiDAR. In contrast, pipeline 5800 relies on both the LiDAR 5806 and another sensor, namely the camera 5802. The operating principles and architectures of pipelines 5700 and 5800 are similar. For example, the preprocessing steps 5708 and 5808 are similar, the plane segmentation steps 5712 and 5812 are similar, the AI steps 5710 and 5810 are similar, and the combinations 5716 and 5816 are similar. Pipeline 5800 similarly combines AI-based keypoint detection with point cloud segmentation to obtain a combined 6DoF pose estimate. The only difference is that Figure 58 the CV / AI pipeline depicted in uses a transformation (T) 5814, which is not a look-up table (LUT) as in the Figure 57 pipeline depicted in. Instead, the transformation or mapping (T) 5814 between the image space of the camera and the LiDAR is obtained through calibration.

[0201] Figure 59 is a schematic diagram of a sensor fusion architecture 5900 for determining the 6DoF pose of a solar panel according to some embodiments. Architecture 5900 combines the Figure 57 and 58 two building blocks shown in. This embodiment combines AI-based keypoint detection for NIR images 5918 from the LiDAR and RGB images from the camera 5908 with point cloud segmentation to arrive at a combined 6DoF pose estimate. The transformation or mapping between the image spaces TNIR (T-nir) 5930 and TRGB (T-rgb) 5940 is obtained accordingly through a LUT or through calibration, as described above. This embodiment allows a heuristic method or meta-model 5904 to determine how to combine (5926) the individual components of the 6DoF pose from different sources (5916, 5928) to form a final 6DoF pose estimate. Depending on the lighting conditions or other environmental information 5902, the meta-model 5904 can use the coordinates (x, y, z) of the panel keypoints based on the camera image 5908 or the LiDAR NIR image 5918 (or some other LiDAR data channel, e.g., Lidar range 5932). For example, the meta-model 5904 can configure the CV / AI pipeline to rely only on the RGB camera 5908 for indoor operation and on the LiDAR reflectivity data channel 6008 for Figure 60), and relies on the LiDAR NIR data channel 5918 for outdoor daytime operation. The meta-model 5904 may also appropriately weight the two sets of key-point coordinates, or even optimize the AI models (e.g., the upper robot AI model 5912, the lower robot AI model 5938 for RGB) for certain environmental conditions (e.g., by selecting appropriate model weights). The meta-model 5904 may also call the post-processing correction 5924 through the lower robot CV / AI pipeline, as described below. According to some embodiments, the preprocessing steps 5910 and 5920 and the plane segmentation step 5934 are as described above. The database 5906 may store data from the environmental sensor 5902 and provide data to the meta-model 5904.

[0202] Figure 60 is a schematic diagram of another sensor fusion architecture 6000 for determining the 6DoF pose of a solar panel according to some embodiments. This example shows a sensor fusion architecture that fuses the traditional analysis of the LiDAR image 6008 with the monocular AI 6004 based on the image from the RGB camera 6002. Three possible methods for calculating the 6DoF pose are shown. One method uses the PnP solver 6006, which takes the pixel positions of the four panel corners (x1,y1),...,(x4,y4) as input; this method can be used indoors. Another method combines the rotation angles (rx,ry,rz) with the world space coordinates (x,y,z) of the panel key points based on the LUT mapping 6014 (from the NIR or reflectivity image space of the LiDAR to the point cloud). This method can be used outdoors. Yet another method is the external calibration 6018 (for the mapping between the image space of the camera and the point cloud). These different methods for calculating the 6DoF pose can be combined in various ways (e.g., for indoor / outdoor). The reflectivity 6008, NIR 6010, and range 6012 can be used for night, daytime, and outdoor, respectively. Other LiDAR data channels (e.g., signal and reflectivity), although not shown, can be additionally or alternatively used for other data channels.

[0203] Figure 61Schematic diagram of a sensor fusion architecture 6100 according to some embodiments. The unrestricted raw images from the NIR LiDAR 6102 are pre-processed (6104) using one of the above techniques and processed by the upper robot AI 6106. The first-stage post-processing 6108 provides a layer of redundancy. This first stage may include the following steps: [1] Determine corners from the NIR image (and / or other channels) in pixel space; [2] Given the panel size, for each corner in pixel space, transform to world space coordinates and calculate the other corners; [3] For each set of four corners in world space, transform back to pixel space; [4] For the four sets (of four) corners in pixel space from the above [3], select a single corner based on some convergence metric (e.g., Euclidean distance); [5] Transform the selected corners from the above (4) in pixel space to world space and its associated set (of other corners); and [6] From the set of corners in the above [5], use a reference corner (i.e., near left). In some embodiments, [7] If the change in the pixel position of the corners in the above [3] exceeds a threshold, use the lower robot vision for a third estimate of 6DoF. The pixel positions of the selected corners are transformed (T) (6124) to obtain the (x, y, z) coordinates of the reference panel corners. This can be combined (6126) with the output of the plane and cylinder segmentation 6122 (which produces an output based on the point cloud output from the LiDAR range 6120) to obtain a first estimate of the 6DoF (X, Y, Z, RX, RY, RZ) of the panel. In the second-stage post-processing 6134, based on the first estimate, the positions of the four panel corners 6128 in coordinates (x, y, z) are further refined based on the two-pass homography algorithm described above. The inverse transformation T -1 (6116) is performed on the positions 6128 to obtain the pixel positions 6110 of the four corners in image space. These pixel positions are input into the two-pass homography algorithm 6112 (an example of which is described above). This includes [8] Determining the first-pass homography matrix H 1 , [9] Determining the second-pass homography matrix H 2 , and

[10] Back-projecting the four panel corners based on the inverse of H 2 . The output of step

[10] is input into the PNP solver 6114 to obtain a second estimate of the 6DoF of the panel. According to [7] described above, this estimate can be input into the lower robot vision (AI) 6130. Figure 61 Explicitly shows the layer of redundancy provided by the first-stage and second-stage post-processing and the lower robot CV / AI pipeline vision. Although Figure 61 not shown, in certain embodiments, the architecture may provide a layer of redundancy including RGB camera images, other LiDAR data channels (e.g., signal and reflectivity), and complementary AI models.

[0204] Range data channels from a LiDAR (e.g., Ouster LiDAR) can be used for depth estimation for a type of segmentation, which is further refined by "two-pass homography" post-processing before being input into a PnP solver. Depth estimation or segmentation can also be performed using a stereo camera, whether in a single-baseline or multi-baseline configuration, as described above.

[0205] Whether using a stereo camera or LiDAR, segmentation of the panel can be achieved by externally calibrating the camera relative to the robotic system (hand-eye). Calibration can also be performed by mapping two image spaces using a homography transformation. For example, in a pure image-based calibration process (i.e., without involving robot removal), the image space of the LiDAR distance image can be mapped to the image space of the camera.

[0206] Alternatively, the 6DoF pose of the panel can be determined solely by the LiDAR. The rotation angles (rx, ry, rz) of the panel in question can be determined by plane segmentation of the point cloud data from the LiDAR. And the (x, y, z) positions of the key points (e.g., corners) of the panel can be determined by mapping the key points found in other data channels (e.g., NIR, reflectivity). To improve the determination of the latter key point positions (x, y, z), the LiDAR can be calibrated relative to a high-resolution camera (e.g., by extrinsic parameters or homography, as described above), such that the key points found in the camera image can be mapped or transformed to the image space of the LiDAR.

[0207] In addition, the various methods of 6DoF pose estimation described above (and their associated post-processing) can be supplemented by additional post-processing (post-processing of the initial 6DoF pose estimation), such as involving different CV / AI pipelines using different sensors (e.g., cameras) attached to the lower robot. For example, the CV / AI pipeline for the lower robot can refine or correct the initial estimate of the 6DoF pose of the panel by profiling the surface contours of the torque tube and / or fixture with structured light (e.g., laser line or grid line projection). The profiling of the torque tube, fixture, or other relevant structures and their known geometries and dimensions can be used as a basis by which the initial 6DoF pose of the panel can be refined. This refinement assumes that the panel mounted on the torque tube is rotationally aligned along the tube, and the midpoint of the fixture coincides with the clamping edge of the mounted panel.

[0208] Figure 62 is a schematic diagram of a solar panel installation 6200 according to some embodiments. The Cartesian axes X, Y, and Z are shown, as well as a panel reference 6202 for six degrees of freedom (6DoF) pose estimation.

[0209] Figure 63FIG. 6300 is a flowchart of a method for autonomous solar installation according to some embodiments. The method includes obtaining (6302) one or more images during installation. The one or more images include images of one or more solar panels and an installation structure.

[0210] The method also includes preprocessing (6304) the one or more images, including compensating for one or more of camera internal parameters or distortion, rectifying the images, and determining depth information. In some embodiments, the preprocessing includes compensating for camera distortion, rectifying the images, and / or determining depth information based on a single-baseline stereo camera, a multi-baseline stereo camera, a time-of-flight sensor, or a LiDAR sensor. According to some embodiments, examples of preprocessing techniques are described above with reference to non-AI methods for pose estimation and exemplary AI models and inferences. The AI model may be trained on images with uncompensated lens distortion and unrectified images.

[0211] The method also includes detecting (6306) the one or more solar panels by inputting the one or more images into one or more neural networks (in series or in parallel) trained to detect solar panels. In some embodiments, the one or more neural networks are trained to output bounding boxes, segmentation, key points, depth, and / or 6DoF pose. According to some embodiments, these techniques are described above in the exemplary AI model and inference section.

[0212] The method also includes a first post-processing (6308) to calculate a first panel pose based on the outputs of the one or more neural networks. The first post-processing may include an AI / CV pipeline different from the AI / CV pipeline used for preprocessing and / or the neural network. The goal may be to further refine the initially estimated panel pose.

[0213] In some embodiments, the first post-processing includes one or more computer vision algorithms (e.g., Ramer-Douglas-Peucker, Hough transform, Shi-Tomasi, homography transform) to process the outputs of the one or more neural networks based on invariant structures in the image (e.g., invariant structures include grid lines and other visual benchmarks on the panel, shadows between panels, projective perspective lines, torque tubes, and relative positions of the panels) to determine the positions of panel key points. According to some embodiments, examples are described above with reference to Figure 51 Described examples.

[0214] In some embodiments, the first post-processing also includes solving Perspective-n-Point based on panel dimensions and panel key points. In some embodiments, the panel key points are the four corners of the panel frame.

[0215] In some embodiments, the method further includes a second post - processing (6310) that includes one or more homography transforms (e.g., the two - pass homography algorithm described above) to obtain a second panel pose of the one or more solar panels based on the first panel pose. The second post - processing compensates for or corrects inaccuracies in the first panel pose based on a visual pattern or fiducial on the solar panel. In some embodiments, the visual pattern includes a grid pattern on the solar panel.

[0216] In some embodiments, the output of the one or more neural networks includes panel segmentation. Figure 52 An exemplary application 5200 of the segmentation according to some embodiments is shown. Figure 55 A schematic diagram of ROI segmentation 5500 according to some embodiments is shown. The second post - processing includes using the panel segmentation as a mask to filter out the background in the image (leaving only the foreground panels) to obtain a masked panel. The second post - processing also includes using a corner - finding algorithm (e.g., Shi - Tomasi algorithm, Harris - Stephens algorithm) to identify the grid intersections in the masked panel as corners. The second post - processing also includes using the corners to calculate a homography matrix H 1 . The second post - processing also includes determining the positions of the grid intersections in millimeter space based on H 1 . The second post - processing also includes calculating a homography matrix H 2 to transform the grid intersections in image space to their positions in millimeter space (a two - dimensional (2D) projection space where the four (physical) corners of the solar panel have the following (x,y) coordinates: (0,0), (0,L), (W,0), (W,L), where W and L are the width and length of the solar panel respectively). The second post - processing also includes back - projecting the corners in millimeter space to image space using the inverse of H 2 .

[0217] In some embodiments, the filtering includes a binary “AND” between the mask and the original image of the solar panel. In some embodiments, determining the positions of the grid intersections includes an association in H 1 space based on the Euclidean distance (e.g., a fraction of the shortest unit size such as the width).

[0218] In some embodiments, the second post - processing includes estimating the four corners k i of the solar panel in image space. The second post - processing also includes calculating a homography matrix H i that maps k 1 to millimeter space. The second post - processing also includes identifying (i) the pixel positions p 1 of the panel features (as fiducials) in image space based on H i ; and (ii) the pixel positions p in millimeter spacei at the corresponding position q i The second post - processing also includes calculating to project p i onto q i of the homography matrix H 2 The second post - processing also includes based on the homography matrix H 2 inverse of (H 2 -1 ) to back - project the corners (0,0), (H,0), (H,W), (0,W) from the millimeter space to the image space.

[0219] In some embodiments, p i and q i include the positions of grid intersections, the centroid of a rhombus, or the common vertex of two triangles (based on grid intersections) and / or pixel structures other than the corners of the solar panel.

[0220] According to some embodiments, various examples of using homography transformation and sensor fusion architecture for 6DoF pose estimation of solar panels are described above with reference to Figures 56 to 61 In particular, the pre - processing and post - processing steps described above with reference to Figures 56 to 61 can be used in the pre - processing, the first post - processing step, and / or the post - processing step of method 6300. For example, Figure 61 illustrates how a two - pass homography algorithm according to some embodiments can be used in post - processing in any sensor fusion architecture based on corner detection.

[0221] In some embodiments, the mounting structure includes a torque tube and a fixture, and the method further includes a third post - processing, including processing one or more images of the torque tube and / or the fixture. In some embodiments, one or more images of the torque tube and / or the fixture are obtained using a high - resolution camera and structured illumination. In some embodiments, the structured illumination is a laser line that is approximately orthogonal or parallel to the torque tube. In some embodiments, the processing of the one or more images of the torque tube and / or the fixture is performed by one or more neural networks and / or computer vision pipelines. In some embodiments, the method further includes locating nuts associated with the fixture by using high - intensity illumination and computer vision algorithms. In some embodiments, the high - intensity illumination is a ring light.

[0222] Returning to reference Figure 63 , the method further includes generating (6312) a control signal based on the first panel pose for operating a robot controller to install the one or more solar panels. In some embodiments, a control signal is further generated (6314) based on the second panel pose. In various embodiments, the control signal is generated based only on the first panel pose, or only on the second panel pose, or the control signal can be generated based on both the first panel pose and the second panel pose.

[0223] Figure 35A A block diagram of an exemplary image processing pipeline 3500 according to some embodiments is shown. The pipeline 3500 includes a module 3502 for acquiring an image, a module 3504 for rectifying the image, a module 3506 for neural network image segmentation of the rectified image, a module 3508 for post-processing the output of module 3506 using computer vision techniques, a module 3510 for performing a Hough transform on the output of module 3508, a module 3512 for filtering and segmenting the Hough lines output by module 3510, a module 3514 for identifying the horizontal and vertical Hough line intersections output by module 3512, a module 3516 for estimating the panel pose based on the horizontal and vertical Hough line intersections (e.g., using 3D panel geometry and the positions of the corners in the image), and a module 3518 for publishing the pose estimate. Figure 35B An exemplary rectified acquired image 3520 (the output of modules 3502 and 3504) is shown, which includes an image of a solar panel 3522 and other objects 3524-2 (e.g., tape) and 3524-4 (e.g., wire). Figure 35C An example is shown according to some embodiments for Figure 35B the exemplary output 3526 of neural network image segmentation for the acquired rectified image shown in Figure 35D An exemplary panel corner detection 3528 (the output of module 3514) according to some embodiments is shown. In this example, corners 3530-2 and 3530-4 are detected based on horizontal lines 3532-4 and 3532-8 and vertical lines 3532-2 and 3532-6.

[0224] Figure 36 Examples 3600 of images 3602, 3606, and 3610 of a road under different lighting conditions and segmentation masks 3604, 3608, and 3612 for these images are shown. Conventional computer vision techniques are useful when the environment is ideal. However, glare, overexposure / underexposure may negatively affect object detection algorithms. Machine learning techniques can overcome environmental inconsistencies, learn from examples with glare and lighting problems, and generalize to new data during inference.

[0225] For non-ideal conditions, non-AI methods can be used, such as using multi-baseline stereovision, LiDAR / ToF. Additionally, non-AI methods can be combined with AI-based methods.

[0226] Exemplary solar panel segmentation

[0227] Some embodiments perform solar panel segmentation by capturing images of solar panels and torque tubes under different lighting conditions.Figure 37A An example of a captured image 3700 including a solar panel 3702 and a torque tube 3704 in accordance with some embodiments is shown. Some embodiments annotate the captured image of the solar panel. Figure 37B An example of an annotated image 3706 (sometimes referred to as an annotated ground truth mask) of the captured image 3700 in accordance with some embodiments is shown. The annotated image includes a black background 3708, an outline of the torque tube 3712 shown in dark gray, and an outline of the solar panel 3710 shown in light gray. Some embodiments create a data set based on the annotated image, use the data set to train an image segmentation model, and use the trained model to detect the solar panel and the torque tube under poor lighting conditions. Figure 37C An exemplary prediction 3714 of the trained model in accordance with some embodiments is shown. The trained model predicts a background 3708, a solar panel 3710, a torque tube 3712, and an object 3716 in the background ( Figure 37A and 37B not shown in

[0228] Some embodiments continuously collect images (and build a data set) and use these images to improve the accuracy of the model. Some embodiments use manual annotation to improve the accuracy of the model. Some embodiments allow a user to adjust parameters of the segmentation model.

[0229] Some embodiments include separate models for semantic segmentation and instance segmentation. Figure 38A An example of image classification is shown. In this example, image classification detects the presence of a bottle 3802, a cup 2806, and a cube 3804. Figure 38B An example of object localization 3816 for the image shown in Figure 38A is shown. In this example, a rectangle 3808 locates the bottle 3802, a rectangle 3810 locates a first cube, a rectangle 3812 locates the cup 3806, and rectangles 3814-2 and 3814-4 locate the cube 3804. Figure 38C An example of semantic segmentation 3818 in accordance with some embodiments is shown. Semantic segmentation helps identify a label 3820 of the bottle 3802, a label 3822 of the cube 3804, and a label 3824 of the cup 3806. Figure 38D An example of instance segmentation 3826 in accordance with some embodiments is shown. In addition to identifying labels 3828 and 3832 of the bottles 3802 and 3806, instance segmentation is capable of distinguishing instances of the cube 3804, thereby determining labels 3830, 3834, and 3836 of the cube 3804 accordingly. Instance segmentation can distinguish multiple solar panel instances in a single image. Instance segmentation generates a mask for each class instance in a camera frame, enabling the identification and localization of individual panels, using the same data collected for semantic segmentation, and supporting illumination invariance.Figure 39 An example of instance segmentation 3900 for a solar panel according to some embodiments is shown. In this example, panel instances 3902, 3904, 3906, 3908, and 3910 are identified. This example shows instances of panels with different orientations (e.g., instances 3902 and 3904).

[0230] Figure 40 An exemplary image processing system 4000 according to some embodiments is shown. System 4000 includes multiple cameras, including camera 4002 for rough positioning, camera 4004 for capturing images when picking up a panel, and camera 4006 for capturing images when placing a panel. Camera 4002 includes a narrow field of view lens, and cameras 4004 and 4006 each include a wide field of view lens. Camera 4002 can be used to identify the trailer position and the initial robot position. In some embodiments, cameras 4004 and 4006 can be the same camera. In some embodiments, camera 4002 can also be used to position the fixture and the central structure during the installation of the solar panel. Cameras 4002, 4004, and 4006 are coupled to corresponding image sensors 4008, 4010, and 4012 (e.g., AR0820 sensors). In some embodiments, the image sensors are optimized for both low light and / or high dynamic range performance. In some embodiments, system 4000 includes a high-speed digital video interface (e.g., FPD-link) and Ethernet for connecting the cameras to one or more GPUs (e.g., GPU 4014 suitable for edge AI processing, such as Nvidia XT, and GPUs suitable for image processing applications, such as Nvidia AGX Xavier TM ). GPU 4016 implements the exemplary image processing pipeline 3500 described above and is connected to the robot controller 4018 using Ethernet. According to some embodiments, in certain systems, GPU4014 can be removed, and the output from the sensor can be directly connected to GPU 4016.

[0231] Some embodiments continue to capture training images when installing solar panels. Figure 41 shows a trailer system 4100 with a rough camera 4102 according to some embodiments, which can be used to capture training images.

[0232] Figure 42A and 42BHistograms 4200 and 4202 of the pose error norms of the neural network with and without using the rough position are correspondingly shown according to some embodiments. As shown, with the rough position, the neural network error (the difference between the true position of the corner of the solar panel and the estimate of the neural network) is significantly reduced (in some cases, from nearly 5 inches to about 0.7 inches).

[0233] Figure 43A and 43B Examples 4300 and 4302 showing the rough positions (e.g., positions 4304, 4036, 4308, and 4310) of the solar panel using the B Mask R-CNN model are shown according to some embodiments. Mask R-CNN is a convolutional neural network (CNN) for image segmentation and instance segmentation. This deep neural network detects objects in an image and generates high-quality segmentation masks for each instance. Mask R-CNN is based on the region-based convolutional neural network. Image segmentation is the process of dividing a digital image into multiple segments or sets of pixels corresponding to the objects in the image. This segmentation is used to locate objects and boundaries (lines, curves, etc.). Mask R-CNN can be used for semantic segmentation and instance segmentation. Semantic segmentation classifies each pixel into a fixed set of classes without distinguishing object instances. In other words, semantic segmentation involves identifying / classifying similar objects at the pixel level as a single class. All objects are classified as a single entity (solar panel). Semantic segmentation is sometimes called background segmentation because it separates the subject of the image (e.g., solar panel, line) from the background. On the other hand, instance segmentation (sometimes called instance recognition) involves correctly detecting all objects in the image while also precisely segmenting each instance. In this sense, instance segmentation combines object detection, object localization, and object classification, and helps distinguish the instances of each object in the image. In addition to having two outputs for each candidate object, including a class label and a bounding box offset, in Mask R-CNN, a third branch also outputs an object mask. This mask output helps extract the finer spatial layout of the object. When compared with other models, in addition to being simpler to train, having good performance, and being efficient, Mask R-CNN is particularly suitable for solar panel recognition because the neural network can perform both semantic segmentation and instance segmentation. In addition, the mask branch only adds a small computational overhead, enabling fast solar panel detection and fast experimentation. Mask R-CNN can be used for image segmentation to identify objects in an image and create a mask within the boundaries of the object.

[0234] Figure 44Illustrated is a system 4400 for solar panel installation according to some embodiments. According to some embodiments, system 4400 includes a main housing 4404, a battery housing 4402, an upper robotic end-of-arm tool (EOAT) 4406, a lower robotic EOAT 4408, and a bracket 4410 for holding a solar panel 4414 on a trailer 4412.

[0235] According to some embodiments, Figure 45A illustrated is a vision system 4502 mounted on a trailer and used to estimate the pose of a structure 4500, and Figure 45B an enlarged view of the vision system 4502 is shown. Various embodiments may mount the vision system on different parts of a ground vehicle, on a robotic arm, or on an end-of-arm tool.

[0236] According to some embodiments, Figure 46A illustrated is a vision system 4602 for module picking 4600, and Figure 46B an enlarged view of the vision system 4602 is shown, which vision system 4602 includes a high-resolution camera with laser line generation.

[0237] Figure 47A Illustrated is a system 4700 for distance measurement at a module angle (i.e., when facing the module) between a position 4702 (an enlarged view of which is shown in Figure 47B and a position 4704 (an enlarged view of which is shown in Figure 47C according to some embodiments.

[0238] Figure 48A Illustrated is a laser line generation system 4800 for detecting the positions of tubes and clamps according to some embodiments. According to some embodiments, Figure 48B an enlarged view of the laser line generation system 4802 is shown, and Figure 48C a view 4804 of laser line generation (horizontal line detecting the clamp, and vertical line detecting the tube) is shown.

[0239] Figure 49A Illustrated is a vision system 4900 for estimating the positions of tubes and clamps and positioning nuts on the clamps according to some embodiments. Figure 49A Also shown is a socket wrench 4902 for tightening the nuts. Figure 49B An enlarged view of the vision system is shown. As Figure 49B shown, the camera uses the laser line described above to locate the tubes and clamps, and uses a flash ring light to locate the nuts on the clamps. The laser provides an accurate estimate of the tube and clamp positions. The flash ring light is used to locate the nuts on the clamps as shown in Figure 49C . When tightened, the nut compresses the clamp to hold the panel in place.

[0240] Exemplary Solar Panel Installation Using Artificial Intelligence

[0241] Figure 50A FIG. 5000 is a flow chart of a method for autonomous solar installation according to some embodiments. The method includes obtaining (5002) an image of an ongoing solar installation. The image includes images of one or more solar panels and one or more torque tubes. In some embodiments, obtaining the image includes using one or more filters to avoid direct solar glare in order to detect an end-of-arm tool (EOAT). In some embodiments, obtaining the image includes using a high-resolution camera with laser line generation to identify the location of one or more torque tubes and / or clamps. In some embodiments, the image includes an image of a clamp and / or a central structure for an ongoing solar installation.

[0242] In some embodiments, the image includes an image of a clamp and / or a central structure for an ongoing solar installation. In some embodiments, a wide-angle fish-eye lens is used to obtain multiple images to create a composite HDR (High Dynamic Range) image within the camera hardware. These images are sent through a Robot Operating System (ROS), which is an advanced software framework for integrating robots and servers, so as to use an OpenCV (image processing framework) module to rectify the images (e.g., change from fish-eye distortion to a planar image). Then, regions and bit depths are selected and used to shrink the HDR image into a standard 8-bit image, effectively cropping the regions and bit depths to use it as an input for training a neural network. At the acquisition point, the robot pose can be stored (using ROS) to create a transformed camera result relative to the trailer (a trailer system for solar panel installation). This can include the robot position and the camera position to identify the position of the image in 3D space.

[0243] The method further includes detecting (5004) a solar panel section by inputting the image into a trained neural network trained to detect solar panels. The neural network can be implemented using software and / or hardware (sometimes referred to as neural network hardware), which uses conventional CPUs, GPUs, ASICs, and / or FPGAs. In some embodiments, the trained neural network includes: (i) a model for semantic segmentation to identify solar panel sections; and (ii) a model for instance segmentation to identify multiple solar panels. In some embodiments, the trained neural network uses the Mask R-CNN framework for segmentation. The trained neural network detects solar panel sections based on features extracted from the image of the ongoing solar installation. In some embodiments, the acquired image is input into the neural network through ROS (e.g., the input image arrives at the neural network module from the OpenCV module). According to some embodiments, reference is made below to Figure 55B describes exemplary techniques for training a neural network. In some embodiments, the neural network performs image segmentation to identify the panel(s) without identifying the position of the panel. In some embodiments, there is a model that implements two functions (semantic segmentation and instance segmentation). Some embodiments use two instances of the same model to optimize throughput. In this case, the camera captures two images, and each image passes through one instance. Running two models allows for processing twice as many images simultaneously.

[0244] The method further includes using a computer vision pipeline to estimate (5006) the panel pose of the one or more solar panels based on the solar panel sections. In some embodiments, the computer vision pipeline includes one or more computer vision algorithms for post - processing, Hough transform, filtering and segmentation of Hough lines, finding horizontal and / or vertical Hough line intersections, and panel pose estimation using a predefined 3D panel geometry and corner positions. In some embodiments, the computer vision pipeline locates fixtures and / or central structures to estimate the panel pose. In some embodiments, the computer vision pipeline locates one or more torque tubes and / or fixture positions to estimate the panel pose. In some embodiments, the computer vision pipeline locates nuts.

[0245] After the nut is located, a socket wrench mounted on a smaller robotic arm can engage with the nut and tighten it to fix the panel in place. Before performing this step, the fixture may become loose and the panel may fall due to the wind.

[0246] In some embodiments, a conventional machine vision hardware is used to perform the panel pose estimation to locate the position of the panel in 3 - D space. In some embodiments, this is a rough identification of circular edges and is not intended to be very precise. Subsequently, the Hough transform can be used to determine the exact position of the edges, which is followed by inferring the edge lines of the panel, determination of the panel intersection positions, and identification of the panel corners. The panel corners are published to identify the position of the panel relative to the robot. For example, based on the panel geometry in 3 - D, the pose of the panel is calculated based on the positions of the corners of the panel in the image.

[0247] In some embodiments, to estimate the panel pose, a computer vision pipeline uses a Perspective-n-Point (PnP) solver with the camera intrinsic parameters (which knows its own camera distortion and parallax). Then, the extrinsic parameters capture the position of the camera relative to the robot using the end effector and EOAT pose at the time of image capture. The robot pose can be continuously captured using timestamps. Then, the timestamps can be used to match the robot pose with the camera acquisition timestamp. In some embodiments, the computer vision pipeline uses the known pose of the robotic arm and end-of-arm tool (where the camera is located) at the time of image capture to calculate the position of one or more corners of the panel.

[0248] The method further includes generating (5008) a control signal based on the estimated panel pose for operating a robot controller to install the one or more solar panels. In some embodiments, after the panel is found, the position is projected along the tube to find the fixture pixels to identify the fixture position (e.g., how far the fixture is, how close it is to the fixture puller). Some embodiments use the fixture position to verify if the fixture is within the allowable window required for the fixture puller on the EOAT. Some embodiments use a central structure to determine the order of placing one or two panels to avoid collision with the fan gear. Some embodiments use the panel position to ensure that the trailer is in an effective position relative to the tube such that the robot is within reach of the work to be performed. Some embodiments use the pose from the leading panel to subsequently guide the lower robot in its fine tube acquisition, which drives the positions of the upper and lower robots for panel placement and nut driving. In some embodiments, the fine tube acquisition described above uses horizontal and vertical lasers to create a profiler system for finding the tube and fixture positions. This refines the working pose from the rough tube from 10 - 20 mm and reduces it to less than plus or minus 5 mm. At the first panel, the rough tube error is within 5 mm, but as this is projected out, the error grows, and the fine tube is used to limit the error to below plus or minus 5 mm.

[0249] Figure 50BFIG. 0 shows a flowchart of a method 5010 for training a neural network for autonomous solar installation according to some embodiments. The method includes: obtaining (5012) a plurality of images of solar panel installations under different lighting conditions; annotating (5014) the plurality of images to identify solar panel images (instead of or in addition to automatically annotating the images, manually annotated images can be used); and training (5016) one or more image segmentation models using the solar panel images to detect solar panels under poor lighting conditions. In some embodiments, various images are used to manually train the neural network, such as different backgrounds (e.g., grassland, dirt) under various weather conditions (e.g., sunny conditions, rainy conditions), different numbers of panels, clamps, and several images of the panels. Within these images, lines are drawn to indicate which pixels represent the panels, clamps, tubes, and central structures. These images and their masks are used to create a series of pseudo-images (sometimes called image augmentation), which are then used by the neural network during the training process. The pseudo-images are input images with angular distortion so that the same input image can be used for training multiple times. For example, 300 - 1000 real (input) images can be used for training, and for each real image, 10 - 20 pseudo-images can be created.

[0250] Embodiments of the present invention have been described above by means of functional building blocks that illustrate specific functions and their relationships. For ease of description, the boundaries of these functional building blocks are arbitrarily defined herein. Alternative boundaries can be defined as long as the specific functions and their relationships are properly executed.

[0251] It will be apparent to those skilled in the art that various modifications and variations can be made to the system for installing solar panels of the present invention without departing from the spirit or scope of the present invention. Accordingly, the present invention is intended to cover modifications and variations of the present invention as long as they fall within the scope of the appended claims and their equivalents. It is to be understood that the language or terms herein are for the purpose of description and not limitation, such that those skilled in the art should interpret the terms or language of this specification according to the teachings and guidance provided.

[0252] The breadth and scope of the present invention should not be limited by any of the above exemplary embodiments, but should be defined only by the following claims and their equivalents.

Claims

1. A method for autonomous solar panel installation, the method comprises: obtaining one or more images during installation, wherein the one or more images include images of one or more solar panels and an installation structure; preprocessing the one or more images, including compensating for camera internal parameters or distortion, rectifying the images, and determining depth information, one or more of which; detecting the one or more solar panels by inputting the one or more images into one or more neural networks trained to detect solar panels; a first post - processing to calculate a first panel pose based on the output of the one or more neural networks; and generating a control signal based on the first panel pose for operating a robot controller to install the one or more solar panels.

2. The method according to claim 1, further comprises: a second post - processing, which includes one or more homography transformations to obtain a second panel pose of the one or more solar panels based on the first panel pose, wherein the second post - processing compensates for or corrects inaccuracies in the first panel pose based on a visual pattern or fiducial on the solar panel, and wherein the control signal for operating the robot controller to install the one or more solar panels is also based on the second panel pose.

3. The method according to claim 2, wherein, the visual pattern includes a grid pattern on the solar panel.

4. The method according to claim 2, wherein, the output of the one or more neural networks includes panel segmentation, and wherein the second post - processing includes: using the panel segmentation as a mask to filter out the background in the image to obtain a masked panel; using a corner - finding algorithm to identify grid intersections in the masked panel as corners; Use the corner to calculate the homography matrix H 1 ; Based on H 1 to determine the positions of grid intersection points in millimeter space; Calculate the homography matrix H 2 , to transform the grid intersection points in the image space to their positions in the millimeter space; and Using the inverse of H 2 back-projects the corners in the millimeter space into the image space.

5. The method according to claim 4, wherein, Determining the position of the grid intersection points includes an association in the H 1 space based on the Euclidean distance.

6. The method according to claim 2, wherein, the output of the one or more neural networks includes panel segmentation, and wherein the second post - processing includes: using the panel segmentation as a mask to filter out the background in the image to obtain a masked panel; using one or more artificial intelligence techniques to identify grid intersections in the masked panel in millimeter space; Calculating the homography matrix H based on the grid intersection points 1 ; and Using the inverse of H 1 back-projects the corners in the millimeter space into the image space.

7. The method according to claim 2, wherein, the second post - processing includes: Estimate the four corners k of the solar panel in the image space i ; Calculate the homography matrix H that maps k i to the millimeter space 1 ; Based on H 1 , identify (i) the pixel position p of the panel feature part in the image space i ; and (ii) the corresponding position q of the pixel position p i in the millimeter space i ; Calculate the homography matrix H that maps p i to q i ; and 2 ; and Based on H 2 -1 , the corners (0, 0), (H, 0), (H, W), and (0, W) are back-projected from the millimeter space to the image space.

8. The method according to any one of claims 1 - 7, wherein, the one or more neural networks are trained to output bounding boxes, segmentation, key points, depth, and / or 6DoF pose.

9. The method according to any one of claims 1 - 8, wherein, the preprocessing includes compensating for camera distortion, rectifying the images, and / or determining depth information based on a single - baseline stereo camera, a multi - baseline stereo camera, a time - of - flight sensor, or a LiDAR sensor.

10. The method according to any one of claims 1 - 9, wherein, the first post - processing includes one or more computer vision algorithms for processing the output of the one or more neural networks based on invariant structures in the image to determine the positions of panel key points.

11. The method according to any one of claims 1-10, wherein, the first post-processing further includes solving Perspective-n-Point based on the panel size and panel key points.

12. The method according to claim 11, wherein, the panel key points are the four corners of the panel frame.

13. The method according to any one of claims 1-12, wherein, the mounting structure includes a torque tube and a fixture, and: wherein the method further includes a third post-processing, and the third post-processing includes processing one or more images of the torque tube and / or the fixture.

14. The method according to claim 13, wherein, the one or more images of the torque tube and / or the fixture are obtained by using a high-resolution camera and structured illumination.

15. The method according to claim 14, wherein, the structured illumination is a laser line that is approximately orthogonal or parallel to the torque tube.

16. The method according to claim 13, wherein, the processing of the one or more images of the torque tube and / or the fixture is performed by one or more neural networks and / or a computer vision pipeline.

17. The method according to any one of claims 13-16 further includes locating nuts associated with the fixture by using high-intensity illumination and computer vision algorithms.

18. The method according to claim 17, wherein, the high-intensity illumination is a ring light.

19. A system for installing a solar panel, the system comprises: a camera system for obtaining one or more images during installation, wherein the one or more images include images of one or more solar panels and a mounting structure; one or more devices for: (i) preprocessing the one or more images, including compensating for camera internal parameters or distortion, rectifying the images, and determining depth information, estimating the panel poses of the one or more solar panels based on solar panel segments; (ii) detecting the one or more solar panels based on the one or more images; and (iii) first post-processing to calculate a first panel pose based on the output of one or more neural networks; and a controller for generating a control signal based on the first panel pose for operating a robot controller to install the one or more solar panels.

20. The system according to claim 19 further comprises: the one or more devices for second post-processing, the second post-processing including one or more homography transforms to obtain a second panel pose of the one or more solar panels based on the first panel pose, wherein the second post-processing compensates for or corrects inaccuracies in the first panel pose based on visual patterns or fiducials on the solar panels, and wherein the control signal for operating the robot controller to install the one or more solar panels is further based on the second panel pose.

Citation Information

Cited By

  • Sunlight environment simulation equipment for detecting service life of photovoltaic module

    CN121036692A

  • A sunlight environment simulation device for lifetime detection of a photovoltaic module

    CN121036692B

  • Pose correction method for mechanical arm of photovoltaic module installation robot and installation robot

    CN121061852A