Lightweight end-to-end defect detection method based on hanging rail type photovoltaic cleaning robot

Through a three-stage detection framework that collaborates with the edge end and the cloud, combined with template matching, color space analysis and visual language model, the problems of high equipment costs, insufficient computing resources and poor robustness of dynamic scenarios are solved, and efficient and safe photovoltaic module defect detection is achieved to adapt to the complex environment of photovoltaic cleaning robots.

CN120411091AActive Publication Date: 2025-08-01HANGZHOU DIANZI UNIV
View PDF 13 Cites 0 Cited by

Patent Information

Application Number
CN202510906038.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-08-01
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

The existing photovoltaic module defect detection methods have problems such as high equipment costs, insufficient computing resources, poor robustness in dynamic scenarios and network instability, making it difficult to adapt to the complex environmental needs of photovoltaic cleaning robots.

Method used

A three-stage collaborative detection framework is adopted, including integrity screening at the edge-end template matching, stain detection at the edge-end color space-based and refined analysis of the cloud-based visual language model, combining dynamic template updates and backup detection mechanisms to achieve lightweight end-to-end defect detection.

Benefits of technology

It improves the reliability and safety of detection, enhances the robustness to environmental interference, reduces the cost of hardware deployment, realizes complementary multimodal advantages, adapts to the needs of complex dynamic scenarios, and promotes the intelligence and automation of photovoltaic power station operation and maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411091A_ABST
    Figure CN120411091A_ABST
Patent Text Reader

Abstract

The invention provides a lightweight end-to-end defect detection method based on a hanging rail type photovoltaic cleaning robot. The photovoltaic cleaning robot collects an original image of the surface of a photovoltaic panel and a position corresponding to the original image; the method comprises the following steps of: after carrying out first round of preprocessing on an original image, extracting a frame line image and a grid line image of a photovoltaic panel in an ROI (Region of Interest) to form a spatial composition profile diagram corresponding to the original image; template matching is carried out, and whether risks exist or not is judged; and according to the network connection condition, performing defect detection by using a visual language model to obtain a first defect detection result, or executing a standby defect detection mechanism by using the edge computing equipment to obtain a second defect detection result. According to the invention, the complex safety detection task in the field cleaning task of the photovoltaic power station can be effectively processed, and the accuracy and reliability of detection are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent operation and maintenance of photovoltaic power stations, and particularly relates to a three-stage collaborative defect detection method for edge computing scenarios such as photovoltaic cleaning robots. This method integrates dynamic template matching based on spatial composition contours, color space analysis, and visual language model (VLM) technology, aiming to achieve a lightweight, highly robust, and highly secure photovoltaic module defect recognition system without relying on environmental perception, and is particularly suitable for photovoltaic panel cleaning operation systems in dynamic and complex scenarios. Background Art

[0002] With the wide spread of photovoltaic power generation, photovoltaic power generation is becoming one of the clean energy sources widely used globally. However, photovoltaic modules are long-term exposed to complex environments (such as high temperature, humidity, dust accumulation, etc.), and surface defects such as hidden cracks and stains are likely to occur, which have a significant impact on the power generation efficiency of photovoltaic modules. Therefore, efficient detection means are urgently needed.

[0003] Traditional photovoltaic module defect detection methods mostly combine various sensor data and modules such as infrared and radar, and are based on computer vision systems, deep learning, and visual language model (VLM) methods to meet the defect detection requirements of photovoltaic power station components. As shown in the following table, relevant enterprises and institutions of higher learning have deeply studied and explored the improvement of defect recognition of photovoltaic modules by combining different detection technologies.

[0004]

[0005] Early solutions relied on infrared thermal imaging (IR) and electroluminescence (EL) imaging technologies to achieve defect detection through hot spot localization and current characteristic analysis. However, such methods require special imaging equipment, have high equipment costs, a single type of detected defect, and are prone to missed detection and false detection. The second-generation solution based on gray histogram segmentation and template matching achieved a response of about 200 ms in a structured scenario, but had a high false positive rate for defects with similar colors (such as bird droppings and delamination with ΔE < 5).

[0006] Existing pure deep learning and visual language models and other models use strategies such as confidence fusion for defect recognition, but have the disadvantage of obstacles to implementation. In particular, although the YOLO series of models (such as the improved YOLOv5) achieved 93.5% mAP in a laboratory environment, their computing power requirements far exceed the carrying capacity of the main control chips of photovoltaic cleaning robots. The actual mAP and frame rate will both drop sharply, resulting in the loss of detection real-time performance and bringing great difficulties to the actual industrial application deployment.

[0007] Existing mainstream detection methods in industrial applications still mainly rely on a single model or a simple combination of models, have not broken through the framework limitations of multi-modal collaboration, have a high degree of manual dependence, high deployment difficulty, and are difficult to adapt to the requirements of complex dynamic scenarios such as photovoltaic cleaning robots.

[0008] The prior art has not effectively integrated the capabilities of rapid screening at the edge and refined analysis in the cloud. Especially when the network is unstable, the traditional solutions cannot guarantee the continuity of basic detection, and there are operation and maintenance safety risks when the cleaning robot is walking.

[0009] In the current industry background of large-scale deployment of photovoltaic power stations and intelligent transformation of operation and maintenance, as a core device to improve power generation efficiency, the photovoltaic cleaning robot urgently needs a detection technology that adapts to the complex environment of photovoltaic power stations and has the safety, real-time performance when the cleaning robot is walking, and can adapt to the limitations of robot deployment. Summary of the Invention

[0010] The present invention aims to solve four core problems in the field of photovoltaic defect detection: 1) Single detection type: Methods based on infrared, EL imaging, etc. have high equipment costs, require installation of multiple sensors, are easily affected by the environment, and are not easy to promote in large-scale photovoltaic power stations; 2) Fast response required at the edge, obstacles in edge deployment, insufficient computing resources, especially the deployment of deep learning architectures is limited; 3) Robustness in dynamic scenarios: Existing template matching technologies are mainly for the recognition and positioning of single target objects (such as fingerprints, numbers, etc.). Due to the motion characteristics of photovoltaic cleaning robots, field of view limitations, and requirements for data processing efficiency in industrial applications, it is difficult to accurately have only a single target board in the field of view. The photovoltaic board scene has large interference and the number of images that can be used to accurately and completely capture the photovoltaic board to be detected is limited, which poses a challenge to the robustness of traditional single-target matching. 4) Network instability: Integrate rapid screening at the edge and refined analysis in the cloud to give play to their respective advantages and ensure the reliable operation of the system in the case of network instability.

[0011] To solve the above technical problems, the present invention provides a lightweight end-to-end defect detection method for photovoltaic power stations. This method is based on the collaborative work of edge computing devices (deployed on photovoltaic cleaning robots) and cloud servers. Its core lies in a three-stage collaborative detection framework, including: integrity screening based on template matching at the edge, stain detection based on color space at the edge (as part of the backup mechanism), and refined defect analysis based on visual language models (VLM) in the cloud.

[0012] In the present invention, a photovoltaic panel array refers to an array composed of several photovoltaic panels of the same size arranged in a line. When the photovoltaic cleaning robot is working, it runs on the surface of the photovoltaic panel array along a straight line parallel to the direction of the photovoltaic panel array.

[0013] The present invention provides a lightweight end-to-end defect detection method based on a rail-mounted photovoltaic cleaning robot, including the following steps:

[0014] Step 1. The photovoltaic cleaning robot runs in a linear cycle on a long strip of photovoltaic panels arranged in a straight line according to the following preset rules: stop, run a distance L1, stop, and run a distance L2;

[0015] The photovoltaic cleaning robot includes an edge computing device and an image acquisition device. The viewing angle of the image acquisition device is adjusted so that the width of the photovoltaic panels in its field of view is not less than the width of the photovoltaic cleaning robot, and the length of the photovoltaic panels in its field of view is not less than the total length L of n photovoltaic panels arranged in a line, and L is not less than the sum of L1 and L2;

[0016] During the stopping, an original image of the photovoltaic panel surface is acquired, and the position of the photovoltaic cleaning robot relative to the photovoltaic panel array is used as the position corresponding to the original image;

[0017] The original image includes a trapezoidal image of photovoltaic panels arranged in a line with a length of L;

[0018] Step 2. After the original image is preprocessed for the first time, a ROI region is obtained by using the four vertices of the trapezoidal image region; the frame line image and grid line image of the photovoltaic panel within the ROI region are extracted to form a spatial composition contour map corresponding to the original image;

[0019] Step 3. On the edge computing device, the spatial composition outline is matched with a template at a corresponding position and a template at a distance L1 from the corresponding position in a preset template library to determine whether there is a risk.

[0020] Step 4. If there is a risk, the photovoltaic cleaning robot stops moving; otherwise, based on the network connection status, the visual language model is used to perform defect detection to obtain a first defect detection result, or the edge computing device is used to execute the backup defect detection mechanism to obtain a second defect detection result.

[0021] Preferably, the determining whether there is a risk comprises the following steps:

[0022] Use the normalized correlation coefficient method to calculate the matching score, and make judgments based on the matching score;

[0023] Match the spatial composition contour map with a template at a corresponding position and a template at a distance L1 from the corresponding position in a preset template library. If the matching score between at least one template and the spatial composition contour map reaches a preset first threshold, it indicates that the photovoltaic panel in the original image is not missing or deformed; otherwise, it indicates that there is a risk.

[0024] The risk mentioned above refers to the risk that the photovoltaic cleaning robot may not be able to clean properly or fall off the photovoltaic panel due to missing or deformed photovoltaic panels;

[0025] The template includes a spatial composition contour template generated based on a trapezoidal image of a number of undamaged and non-deformed photovoltaic panels arranged in a single-row pattern with a length of L.

[0026] Preferably, the calculation process of the matching degree score includes the following steps:

[0027] Calculate the normalized cross-correlation coefficient R(x, y) between the spatial composition contour map I and each template T according to the following formula:

[0028]

[0029]

[0030]

[0031] where (x, y) is the reference coordinate of the starting point of the current match in the spatial composition contour map I, (i, j) is the offset relative to the reference coordinate, with the value range of i ∈ (0, w−1) and j ∈ (0, h−1), where w and h are the width and height dimensions of the template T respectively, and I(x, y) represents the pixel value of a point in the sub-block area with the same size as the template image T with (x, y) as the upper left corner point in the spatial composition contour map I; μ I is the gray mean value of the current sub-block area, and μ T is the gray mean value of the template image T. I(x+i, y+j) represents the pixel value of a point in a sub-block with the same size as T o in I, with its upper left corner located at (x, y) and the lower right corner located at (x+w-1, y+h-1), and this sub-block will calculate the similarity with T o pixel by pixel. T(i, j) corresponds to I(x+i, y+j) one by one, representing the pixel value at the offset (i, j) in the template image; the value range of R(x, y) is [-1,1], and the closer the value is to 1, the higher the matching degree.

[0032] Preferably, step 3 further includes the following steps:

[0033] When the matching score is greater than the first preset threshold, extract the matching area with the highest matching score, and calculate the Mahalanobis distance D between this area and the corresponding template; if D is less than the preset second threshold, use the weighted moving average method to fuse and update the pixel value, mean value and variance of the corresponding template; if D is not less than the preset second threshold, retain the old template, and this image is marked as having no missing parts but containing transient interference.

[0034] Preferably, a polarizer is provided on the image acquisition device; the first-round preprocessing includes light compensation and noise filtering. The image acquisition, first-round preprocessing, and standby defect detection mechanism at the time when the cleaning robot stopped last time are executed in parallel with the image acquisition at the next stop, and a double-buffer queue is used to manage the image data to be processed, thereby reducing the waiting time between tasks.

[0035] Preferably, in step 4, the defect detection is performed using a vision language model according to the network connection situation to obtain a first defect detection result, or a standby defect detection mechanism is executed by an edge computing device to obtain a second defect detection result, which specifically includes:

[0036] When the network connection between the edge computing device and the cloud server is normal, the images determined to be risk-free are input into the vision language model in the cloud, and defect detection is performed on the images in combination with the photovoltaic defect semantic information to obtain a first defect detection result;

[0037] When the network connection is abnormal, the images determined to be risk-free are cached on the edge computing device, and a standby detection mechanism is executed. The standby detection mechanism includes stain defect detection based on a color space to obtain a second defect detection result; and after the network connection returns to normal, the cached images are uploaded to the vision language model in the cloud to obtain a first defect detection result, and the first defect detection result and the second defect detection result are fused to obtain a final defect detection result.

[0038] Preferably, before the images determined to be risk-free are input into the vision language model in the cloud, the following steps are further included: performing a second-round preprocessing on the images determined to be risk-free, and the second-round preprocessing includes inverse perspective transformation. Inverse perspective transformation changes the picture from a trapezoid back to a normal shape, which is more convenient for the defect recognition of the VLM.

[0039] Preferably, the fusion of the first defect detection result and the second defect detection result includes the following steps: the final defect detection result includes two types of defects, namely stains and glass breakage; the stain category further includes bird droppings and dirt; it is judged whether the first defect detection result is glass breakage; if so, the final defect detection result is glass breakage, otherwise: it is judged whether the VLM detects stains; if not, the second defect detection result is adopted; it is judged whether the confidence level of the stains detected by the VLM is less than a third threshold; if so, the second defect detection result is adopted; if not, the first defect detection result is adopted. The above design considers that the edge-side standby mechanism (the second defect detection result) is mainly based on the color space and is more sensitive to "stain" type defects (such as bird droppings and dirt). For "glass breakage" type defects, the color space standby mechanism at the edge has a weak detection ability for "glass breakage" and mainly relies on the first defect detection result.

[0040] Preferably, in step 1, n is 3, 4 or 5; with the values of the above sizes, the trapezoidal pictures obtained have appropriate sizes, which can reflect common defects such as bird droppings, dirt, and glass breakage, and also meet the composition of the spatial composition contour map. When n is too small, such as 1, there is no possibility of forming a spatial composition contour map. When the value of n is too large, the proportion of each photovoltaic panel in the formed spatial composition contour map is too small, and it is not easy to detect the defects on it;

[0041] In step 2, the border line image and grid line image of the photovoltaic panel in the ROI area are extracted, including: by processing the image of the ROI area through an edge detection algorithm and a Hough line segment detection algorithm, the border line image and grid line image of the photovoltaic panel in the ROI area are obtained. Through the above two edge detection algorithms and Hough line segment detection algorithms, the extracted border line image and grid line image are clearer, which is beneficial for later template comparison and risk identification.

[0042] Preferably, the vision-language model is the Qwen-VL model, and it is fine-tuned using a dataset containing photovoltaic module images and their corresponding defect descriptions.

[0043] The defects include glass breakage, bird droppings, and dirt;

[0044] The prompt word framework of the vision-language model includes the following dimensions:

[0045] Spatial dimension: The relative position description of the defect on the component;

[0046] Morphological dimension: The crack trend or stain diffusion morphological characteristics;

[0047] Temporal dimension: Comparative analysis with historical detection results.

[0048] The beneficial effects of the present invention:

[0049] 1. Significantly improve the reliability and safety of structural risk detection: Based on the matching strategy of "spatial composition contour" and the dynamic template update mechanism, it can more accurately identify high-risk risks such as missing photovoltaic panels and severe deformation, effectively ensuring the safe operation of equipment such as cleaning robots.

[0050] 2. Enhance the robustness to environmental gradual changes and instantaneous interferences: The scene-based dynamic template library combined with GPS and the update mechanism based on Mahalanobis distance enable the system to adapt to the progressive changes of photovoltaic panels due to aging, dust accumulation, etc., and effectively resist instantaneous interferences such as sudden illumination changes and short-term occlusions, ensuring the long-term stability and accuracy of detection.

[0051] 3. Achieved lightweight deployment and efficient collaboration: By placing computationally intensive tasks (such as VLM analysis) in the cloud and performing fast integrity screening, dynamic template management, and backup detection at the edge, the overall solution is made lightweight, suitable for embedded platforms with limited computing power, and reduces the hardware deployment cost.

[0052] 4. Complementary advantages of multi-modalities to enhance comprehensive detection ability: It integrates template matching (sensitive to structural defects and deformations), color space analysis (sensitive to specific stains), and VLM (powerful graph-text understanding and fine-grained recognition capabilities), overcomes the limitations of single models, and improves the comprehensive detection ability for various types of defects.

[0053] 5. Has good engineering characteristics and application potential: The entire method has a clear process, and the modular design is easy to deploy and integrate. It can adapt to the complex dynamic scene requirements in the actual industrial environment, promoting the intelligence and automation level of photovoltaic power station operation and maintenance. Brief Description of the Drawings

[0054] Figure 1 is the process schematic diagram of the present invention;

[0055] Figure 2 is the process schematic diagram of the VLM model of the present invention;

[0056] Figure 3 is the walking schematic diagram of the photovoltaic cleaning robot in the present invention;

[0057] Figure 4 is the defect detection effect diagram of different types of stain categories in the backup detection mechanism of the present invention;

[0058] Figure 5 is an example diagram of the preprocessed process and the ROI area template image of the on-site photovoltaic panel. Detailed Embodiment

[0059] The lightweight end-to-end defect detection method based on the rail-mounted photovoltaic cleaning robot provided by the present invention, as Figure 1 shown in the process schematic diagram, specifically includes the following steps:

[0060] S1. Image acquisition: The image acquisition device (such as a camera) deployed on the photovoltaic cleaning robot captures the image data of the photovoltaic module and obtains the position of the photovoltaic cleaning robot relative to the photovoltaic panel array as the position corresponding to the original image.

[0061] S2. Image preprocessing: In this stage, multiple processes are performed on the original image to provide high-quality small images for subsequent template matching and defect analysis.

[0062] Lighting compensation and noise filtering (the first round of preprocessing):

[0063] 1) Light compensation: For the images of the defective scenes to be judged, the Contrast Limited Adaptive Histogram Equalization (CLAHE) algorithm is used to adjust the local contrast of the images and reduce the impact of uneven illumination.

[0064] 2) Noise filtering: The Non-local Means denoising algorithm is adopted to effectively filter out the image noise while retaining the edge and texture details.

[0065] 3) Defining the region of interest, i.e., the ROI region: Based on the actual situation of the photovoltaic power station, the four vertices of the trapezoidal image region of the photovoltaic panels arranged in a straight line with a length of L in the original image are used to obtain the ROI region. This ROI definition method takes into account the actual shooting requirements of the photovoltaic cleaning robot. The border line image and grid line image of the photovoltaic panels within the ROI region are extracted to form a spatial composition contour map corresponding to the original image, such as Figure 5 the ROI definition in

[0066] Inverse perspective deformation correction (the second round of preprocessing, after the integrity screening): Before the images determined to be risk-free are input into the visual language model in the cloud, the perspective images are corrected into top views without perspective effects using the internal and external camera parameters. The innovation lies in clearly stating that this step is carried out after the preliminary component integrity screening, avoiding the problem of preprocessing failure caused by the inability to extract feature points due to component missing or deformation.

[0067] S3, Edge side - Fast screening and tracking of integrity based on template matching of spatial position composition

[0068] Abandon the traditional single-object template matching and instead match whether the "spatial composition" formed by the contours of several photovoltaic panels within the current position ROI coincides with the composition in the template library at this position. This is one of the key innovations of this patent. When the target photovoltaic panel is missing or severely deformed, its "absence" itself becomes a part of the characteristics of the spatial composition, enabling a more accurate judgment of "absence" or "severe structural risk". Due to the limited field of view during the movement of the robot, the opportunity to accurately photograph a single panel is limited, and multi-panel composition matching is more necessary.

[0069] "Component missing and deformation" means that during the walking process of the cleaning robot on the photovoltaic panels, it may encounter structural factors such as partial missing of the photovoltaic panels, deformation of the photovoltaic panels, and tilting of the photovoltaic panels, which affect the walking of the robot.

[0070] For a PV panel confirmed to have no missing deformation, when the robot walks for the first time, it collects and generates a negative sample image of "no missing deformation" and its GPS latitude and longitude position (Latitude, Longitude) information, which constitutes the initial template library. After the first round of preprocessing of the image, the spatial composition contour template is extracted. When performing template matching, the system will select the template image corresponding to the current shooting position in the template library and the template at a position with a distance L1 from the corresponding position to form the candidate template library for this matching.

[0071] On the edge computing device, use the template matching algorithm preset for the spatial composition contour of the PV module to analyze the spatial composition contour map of the image after the first round of preprocessing. Calculate the matching degree score between the current image and the candidate template library, and set a first preset threshold.

[0072] After the matching degree score calculated using the normalized correlation coefficient method is higher than the first threshold, further calculate the Mahalanobis distance D between the current matching area and the template image. Only when the Mahalanobis distance is less than the preset second threshold is the matching considered valid and there is no large instantaneous interference, and only then is the template updated. The template pixels are updated using the weighted sliding average method. At the same time, the pixel-level mean and variance of the template are also dynamically updated, enabling the template to track the gradual changes of the target while resisting instantaneous interference and avoiding template contamination. If D is greater than the preset second threshold, the old template is retained, and this image is marked as having no missing but containing instantaneous interference.

[0073] S4. Edge - Emergency Risk Judgment and Preliminary Shunting

[0074] Make a judgment based on the matching degree score in step S3:

[0075] If the condition of "the matching scores of at least one template and the spatial composition contour map in two consecutive frames reach the preset first threshold" is not met, it is determined that there is a risk in the current detection area. The risk refers to the risk that the PV cleaning robot cannot clean normally or falls from the PV panel due to missing or deformed PV panels. The edge device immediately triggers the robot stop mechanism and reports the risk status, and the current detection process terminates to ensure the safety of the robot.

[0076] If the matching scores of at least one template and the spatial composition contour map in two consecutive frames reach a preset first threshold, it means that there is no missing deformation of the photovoltaic panel in the original image, and it is determined that there is no obvious risk of missing deformation in the current detection area. According to the network connection situation, use the visual language model to perform defect detection to obtain the first defect detection result, or use the edge computing device to execute the backup defect detection mechanism to obtain the second defect detection result. At the same time, the cleaning robot can continue to move along the predetermined path to collect images of the next photovoltaic module. At this time, if there is a matching score greater than the first preset threshold, extract the matching area with the highest matching score, and calculate the Mahalanobis distance D between this area and the corresponding template; if D is less than the preset second threshold, use the weighted moving average method to fuse and update the pixel values, mean, and variance of the corresponding template; if D is greater than the preset second threshold, retain the old template, and this image is marked as having no missing but containing instantaneous interference.

[0077] S5, Cloud - VLM Refined Defect Detection (Main Path), as Figure 2 Schematic diagram of the VLM model process;

[0078] Judge the network connection status between the edge device and the cloud server:

[0079] If the network connection is normal: Upload the pre - processed image to the cloud server. The visual language model (VLM) deployed in the cloud receives the image and combines the preset photovoltaic defect semantic information (input through optimized prompts, and the prompt framework includes spatial dimension, morphological dimension, and temporal dimension) to perform detailed defect detection and classification on the image (identifying two types of defects: glass breakage and stains). The VLM directly outputs the detailed defect recognition result as the first defect detection result.

[0080] If the network connection is interrupted or unstable: The image cannot be immediately uploaded to the cloud. At this time, cache the image determined to be risk - free in the local storage of the edge device and wait for the network to recover before re - attempting to upload it to the cloud VLM for analysis. At the same time, to avoid interrupting the basic defect detection, start the edge - side backup detection mechanism (Step S6).

[0081] S6, Edge - side - Backup Detection Mechanism and Result Fusion (Backup Path)

[0082] When the network is interrupted or as a supplementary verification, execute the following backup detection process on the edge device:

[0083] 6a. Stain defect detection: Convert the image to the HSV (Hue, Saturation, Value) color space, and use a multi-channel adaptive segmentation algorithm to distinguish the surface of a normal photovoltaic panel from areas with attachments such as stains and bird droppings based on features such as color and saturation, obtaining the second defect detection result. Before the network connection is restored, the detection result of the photovoltaic panel is output according to the result of this algorithm.

[0084] 6b. Result caching and decision fusion: Cache the detection result of step 6a. After the network connection is restored, the locally cached image will be successfully uploaded to the cloud VLM for analysis (performed according to step S5), obtaining the first defect detection result. After obtaining the detection result of the VLM, fuse the first defect detection result and the second defect detection result based on the third confidence level to obtain the final defect detection result. This fusion strategy ensures that preliminary detection can be performed even when the network is poor, and uses the high-precision analysis ability of the VLM to confirm and correct after the network is restored.

[0085] Preferably, in S2, template matching (S3) does not need to rely on the feature points at the four corners of the component for inverse perspective correction, because if the component is missing or deformed, the feature points may not be extracted, resulting in the failure of preprocessing. Therefore, light compensation and noise filtering are used as the first round of preprocessing. After the detection of missing and deformed components without safety risks, perspective transformation is used as the second round of image preprocessing and input for the next analysis (S5, S6).

[0086] Furthermore, in S3, since it is difficult for the camera to accurately and separately capture the target photovoltaic panel, the collected original image will include the target photovoltaic panel and 2-4 adjacent photovoltaic panels. When the target photovoltaic panel is missing or deformed, ordinary template matching may also determine that there is no missing or deformed situation because it matches the integrity of the adjacent photovoltaic panels. Under the ideal condition of assuming that the contour structure of the photovoltaic panel is complete, there are only spatial composition differences in the images of 3-5 photovoltaic panels collected by the camera at different positions. Therefore, considering the actual situation of the photovoltaic power station, we change the core of template matching detection to whether the spatial composition of the photovoltaic panel contour in the current position ROI area matches the template library of this position.

[0087] Based on this, the core of the template matching algorithm preset for the spatial composition contour of photovoltaic components is the complete matching of the contour structure of the target photovoltaic panel and the photovoltaic panels in adjacent positions in the matching contour template library, without the need for additional sensor data or target detection means. Its core structure includes:

[0088] 1. Data preparation and preprocessing module: Responsible for collecting and screening templates / test samples in different regions, and performing the first round of preprocessing and template size normalization.

[0089] 2. Template Library Generation Module: The robot is equipped with a camera to collect negative sample images of the first row of photovoltaic panels in the photovoltaic power station and obtain the latitude and longitude positions (Latitude, Longitude) of the photovoltaic panels. After the first round of preprocessing, "non-missing and deformed" negative sample images are generated, and through light compensation, noise filtering, edge detection, Hough line segment detection, and ROI delineation, a set of spatial composition contour templates (such as Figure 5 - Spatial composition contour templates) are generated to form the template library {T pre,neg_1 ,T pre,neg_2 ,..., T pre,neg_N} of the first row of photovoltaic panels.

[0090] 3. Test Image Processing Module: Perform the same preprocessing on the input image to be detected as on the training samples.

[0091] 4. Template Matching Calculation Module:

[0092] The GPS positioning module of the robot provides a location information, transmitting the latitude and longitude positions (Latitude, Longitude) of the current image capture. After that, when the robot conducts the first round of walking and shooting, it selects the three template images after the position (Latitude, Longitude) for matching. Since the second round, each time template matching is performed, the template corresponding to the current shooting position (Latitude, Longitude) in the template library and the template at the position with a distance L1 from the corresponding position are selected to form the template library for this matching. This is the core embodiment based on the spatial position composition.

[0093] In the template matching stage, the normalized cross-correlation algorithm (NCC) is adopted to calculate the matching scores between the spatial composition contours of the photovoltaic panels in the test image and the spatial composition contours of each template in the template library.

[0094] 5. Decision Making and Dynamic Template Fusion Module:

[0095] To solve the problem of matching failure caused by various changes during long-term operation in traditional static template matching, the present invention introduces a mechanism that combines template matching based on spatial composition contours with dynamic template library update.

[0096] The matching scores calculated based on each template image are compared with a preset first threshold. When there is a matching score > the first threshold, the matching region with the highest matching score is extracted, and the Mahalanobis distance D between the current region and the template image is calculated and compared with a preset second threshold.

[0097] Template update condition: If D < the preset second threshold, it is considered that the matching is valid and there is no large instantaneous interference. At this time, the mean, variance, and template are dynamically updated (Pcurrent For the current matching area), this image is marked as non - missing. If D > the preset second threshold, the old template is retained, and this image is marked as non - missing but containing transient interference. If the condition of "for at least one template in two consecutive frames, the matching score with the spatial composition contour map reaches the preset first threshold" is not met, it is determined that there is a risk and an alarm is triggered.

[0098] Preferably, the VLM model selects a large - scale vision - language model of the Qwen - VL series (such as Qwen - VL - Max). To improve its performance in the photovoltaic defect detection task, instruction fine - tuning is performed based on photovoltaic defect text - image pairs to optimize cross - modal alignment. The processing flow includes:

[0099] 1. Instruction fine - tuning: Use a dataset containing photovoltaic component images and their corresponding defect descriptions (text) to fine - tune the pre - trained VLM, enhancing the model's understanding of photovoltaic - specific defect scenarios and cross - modal alignment ability.

[0100] 2. Standardized API interface design: Design a standardized API interface for data transmission (image upload, instruction issuance) and result parsing (defect type, location, confidence, etc.) between edge devices and the cloud - based VLM to ensure system integration.

[0101] 3. Multi - image comparison mechanism (optional): When the VLM analyzes, optionally load cloud - based historical images or standard reference images related to the component to be detected. By comparing the similarity with the input image, it helps to determine whether the defect is a newly added, developing defect, or helps to confirm the same - type defect pattern, improving the diagnostic accuracy. This mechanism can be flexibly configured to be enabled or not.

[0102] 4. Dynamic prompt injection: Use an attention - guided prompt - word optimization mechanism to dynamically inject a structured photovoltaic defect semantic template (for example, a three - dimensional text description framework preset with dimensions such as spatial location, morphological features, and possible causes) into the input prompt of the VLM. This helps to guide the VLM's attention to specific defect features, compensate for the blind spots in edge - side detection, and improve the recognition accuracy of complex or subtle defects (such as specific types of glass fractures).

[0103] Preferably, the photovoltaic defect semantic template library can contain multi - dimensional information. For example: Spatial - dimension prompt: Such as "Check whether there are linear cracks in the upper - left corner edge area of the photovoltaic component"; Morphological - dimension prompt: Such as "Find whether there is a fragmented area with a white spider - web - like radiation morphology and blurred boundaries"; Context / temporal - dimension prompt (if there is historical data): Such as "Compare with historical images to confirm whether the previously marked stain area has expanded".

[0104] Preferably, to ensure the real-time performance of the system, an asynchronous pipeline architecture can be adopted during deployment: tasks such as image preprocessing (step S2), edge-side detection (steps S3, S6a), etc. are executed in parallel with image acquisition (step S1). A double-buffer queue or a similar mechanism is used to manage the image data to be processed, reducing the waiting time between tasks and achieving near-zero waiting data flow. If the edge device hardware supports it (such as including an NPU), some edge-side model inference tasks (such as the computationally intensive parts of template matching or color segmentation) can be assigned to dedicated hardware cores for execution to further accelerate the processing. The VLM inference on the cloud can also utilize a GPU or a dedicated AI accelerator.

[0105] Preferably, the result fusion strategy (step S6b) additionally incorporates dynamic fusion of confidence levels. Requirements for the confidence level of the VLM output are added to the VLM prompt words, and the problem template is modified so that the VLM outputs the confidence level of its image analysis. When detecting "stain" type defects, when the confidence level output by the VLM < the third threshold (usually set to 0.7 during training), the results of the backup detection mechanism are adopted; otherwise, the detection results of the VLM model are output, avoiding the subjectivity of semantic analysis. When detecting "glass breakage" type defects, since the detection ability of the color space backup mechanism is weak, the results of the VLM model are completely adopted.

[0106] To better describe the working principle and details of the present invention, the following further describes the embodiments of the present invention in conjunction with the accompanying drawings to more clearly clarify its objectives, technical solutions, and positive effects.

[0107] The image acquisition device (for example, a camera of model HX2-TC2) deployed on the photovoltaic cleaning robot has an additional polarizer added in front of its lens to capture the image data of the photovoltaic module. When the photovoltaic cleaning robot is working, it runs along a straight line parallel to the direction of the photovoltaic panel array on the surface of the photovoltaic panel array (for example, grasping the guide rails on the left and right sides of the photovoltaic panel array), as Figure 3 shown. The normal walking speed of the robot is set to 18 m / min. Its shooting logic is as follows:

[0108] Fixed-point shooting sequence: Subsequently, the robot stops moving every time it advances approximately 30 centimeters (i.e., every 1 second, calculated based on a speed of 18 m / min ≈ 30 cm / s) and performs fixed-point shooting. At each fixed-point position, the camera takes 1 image.

[0109] Advancing and repeated shooting: After completing two fixed-point shootings (2 images), the robot continues to walk forward 1.4 meters without shooting. Then it enters the above-mentioned fixed-point shooting sequence again (i.e., stops every 30 centimeters and takes 1 image).

[0110] Image usage: One image captured at each fixed point will be used as the original image data for detecting defects such as deformation and missing of the next or current area of photovoltaic panels in subsequent steps (especially in S3).

[0111] This specific walking and shooting logic ensures that images can be obtained at regular intervals and with redundancy (one image is captured at each point) during the robot's movement, while also considering the operation efficiency (by not shooting during movement at a specific distance).

[0112] The following first-round preprocessing operations of light compensation and noise filtering are performed on the obtained images to improve the image quality and eliminate environmental interference:

[0113] Light compensation processing (such as Figure 5 - Light compensation): The contrast-limited adaptive histogram equalization (CLAHE) algorithm is used to adjust the local contrast of the image, improve the detail performance, and enhance the recognizability of defects.

[0114] Noise filtering (such as Figure 5 - NLM denoising): The non-local means (NLM) denoising algorithm is used to remove the random noise and environmental interference in the image acquisition while ensuring the integrity of edges and texture details.

[0115] Traditional single-object template matching is prone to failure when the target photovoltaic panel is missing or severely deformed. The "spatial composition contour" matching strategy adopted in the present invention uses the contour arrangement characteristics (i.e., "spatial composition") formed by the target photovoltaic panel and its adjacent 2 - 4 photovoltaic panels within the ROI as the matching object. Even if the target photovoltaic panel is missing, the "missing" state itself will change the characteristics of this spatial composition, making this "missing" recognizable, rather than resulting in a matching failure or mis-matching to the intact photovoltaic panel next to it. Considering that during the movement of the cleaning robot, the field of view is limited, and the two sides of the photovoltaic panel are usually overhead, it is physically difficult to stably achieve accurately shooting only a single target panel, and there is often only one chance to accurately shoot the target panel. Therefore, matching based on the "spatial composition" containing the contour information of the surrounding environment (adjacent photovoltaic panels) is a necessary means to improve the detection robustness and accuracy.

[0116] The edge component integrity is quickly screened for the spatial composition contour image after the first-round preprocessing, and a preset component spatial composition contour template matching algorithm is deployed on the edge computing device. The matching degree between the spatial composition contour image after the first-round preprocessing and the template is calculated to obtain a matching score. A first threshold is predefined to judge the missing risk, and when the matching degree < this threshold, it is determined as abnormal.

[0117] To address the defects of traditional static template matching, such as the position offset of the target photovoltaic panel in the test image caused by light changes, manual target adjustment, camera parameter changes, and errors in the shooting time at a fixed point during long-term operation, a combination of template matching based on spatial composition contours and dynamic template library update is invented. The actual camera is mounted on a cleaning robot, with the viewing angle inclination and installation angle fixed, without the interference of common rotation invariance problems. And through learning at the edge, the template can track the gradual changes of the target while resisting instantaneous interference (such as short-term occlusion, noise).

[0118] The specific detection process is as follows:

[0119] When the robot performs the first walking operation in a specific area of the photovoltaic array (such as the first row of photovoltaic panels), it will collect the reference images of "no missing deformation" at each position along the way through the mounted camera. At the same time, the accurate latitude and longitude positions (Latitude, Longitude) corresponding to these images are obtained through the GPS positioning module. The collected original images are first preprocessed in the first round in S2, including light compensation and noise filtering. After ROI delineation, the ROI image is obtained. Subsequently, edge detection (Canny operator or other edge detection algorithms) and Hough line detection are performed on the ROI image to extract the contour information of the photovoltaic panel (and its adjacent panels), and it can be converted into a binary spatial composition contour template (such as the spatial composition contour template shown Figure 5 . These template images containing spatial composition contour information, together with their corresponding GPS position information, as well as the initial pixel-level mean and variance statistics for subsequent Mahalanobis distance calculation, are stored together to form the initial template library {T pre,neg_1, T pre,neg_2 ,..., T pre,neg_N} at this fixed point position.

[0120] In the subsequent regular detection process, when the cleaning robot moves to a new position and collects an image, the system will first obtain the current GPS positioning information (Latitude, Longitude) of the robot. Then, according to this position information, the template image at the current shooting position (Latitude, Longitude) and the template at a position with a distance L1 from its position in the template library are selected from the pre-stored template library to jointly form the candidate template library for this matching operation. This way of selecting candidate templates based on GPS position reflects the core idea of "spatial position composition", making the matching more targeted and capable of adapting to the subtle light or photovoltaic panel arrangement differences that may exist in different geographical locations.

[0121] For the current frame image to be detected, the first round of preprocessing in S2, which is consistent with the template library sample generation, is first performed to extract its spatial composition contour information within the ROI. Then, the normalized cross-correlation algorithm is used to calculate the matching score R(x, y) between the spatial composition contour of the ROI of the image to be detected and the spatial composition contour of each template in the candidate template set. The normalized correlation coefficient formula is as follows:

[0122]

[0123]

[0124]

[0125] (x, y): The reference coordinates of the current match in the spatial composition contour image (e.g., the upper left corner of the photovoltaic panel area), indicating the starting point of the current match. (i, j): The offset relative to the reference coordinates (traversing the entire template size), with a value range of i∈(0, w−1), j∈(0, h−1), where w×h is T o I (x, y) means the top left corner of I and T o μ is the pixel value of a point in a sub-block area of the same size. I is the grayscale mean of the sub-block in the current spatial composition contour image, μ T T o The grayscale mean of . I(x+i, y+j) represents the grayscale mean of one of I and T o The pixel value of a point in a sub-block of the same size, with its upper left corner at (x, y) and its lower right corner at (x+w-1, y+h-1), will be the same as T o Calculate similarity pixel by pixel. T(i, j) corresponds one-to-one to I(x+i, y+j), representing the pixel value at offset (i, j) in the template image.

[0126] The value range of R(x, y) is [-1,1]. The closer the value is to 1, the higher the matching degree.

[0127] For each template in the candidate template group, the calculated NCC matching score is compared with a preset first threshold (for example, usually set to 0.7). If among all the candidate templates, there is at least one template with the highest NCC matching score greater than the first threshold, then the preliminary match is considered successful, and the spatial composition contour information I in the current frame ROI is extracted. current and its corresponding original template T oldIf the highest NCC matching scores of all candidate templates are less than the first threshold, it means that the current image to be detected has a low matching degree with all templates representing "no missing deformation" samples. Therefore, it is determined that there may be a missing photovoltaic module or a serious structural risk in the current area, and the subsequent S4 alarm process is triggered.

[0128] For the initially successfully matched region I current and template T old , the weighted moving average method is adopted to retain slow-changing factors such as light, dust, and the aging of photovoltaic panels. Also, considering the instantaneous and short-term occlusion interference problems such as flying birds and local reflections that may exist in the field, the pixel-level mean and variance of the maintenance template are added, combined with the Mahalanobis distance constraint. The specific formula is as follows:

[0129]

[0130] In the matching stage, calculate the Mahalanobis distance between the spatial composition contour image and the template image. Only when D(x, y) < the second threshold (usually set to 3.0 based on training data), the matching is considered valid and template update is allowed.

[0131]

[0132] The weighted moving average formula, where the parameter α ∈ [0, 1] is the forgetting factor, generally chosen as 0.9 (slow update, suitable for scenarios of gradual light change and mechanical wear, avoiding sudden interference), T old is the pixel value of the old template at the position (x, y) (similarly for T new ), I current is the pixel value of the matching region in the current spatial composition contour image, with the position (x, y). T_new is neither T_old nor I current , but the new template fused by the two through the Mahalanobis distance, which will replace T_old and be stored back in the template library at the corresponding GPS position.

[0133]

[0134]

[0135] The mean and variance are the core parameters of the Mahalanobis distance, dynamically evaluating the difference between the current data and the template and deciding whether to update the position (Latitude, Longitude) template. The mean update is consistent with the weighted average, reflecting the long-term trend. The variance update consists of two parts: 1. The local variance σ 2 current of the current window (measuring the noise intensity of the current area). 2. The squared difference between the new and old means (μ current − μ old ) 2(Measuring the amplitude of gradual changes). In this way, the template's judgment criteria for "what is normal change" are also dynamically adjusted, making it more adaptable to the gradual changes in the environment and the aging of components.

[0136] When D(x,y) ≧ the second threshold, it indicates that although the matching score is high, the current matching region I current and the template T old have significant differences in statistical characteristics, and there may be large transient interferences or atypical, non-structural changes. To prevent these abnormal situations from contaminating the purity of the template, the system chooses not to update the template at this time and retains the old template T old . And mark the current image as "no missing but containing transient interference". The mechanism of this set, which includes preliminary matching, Mahalanobis distance secondary verification, and dynamic update, has the following core advantages: 1) Introduce the Mahalanobis distance as the key constraint condition for template update and fusion, rather than simply using the Mahalanobis distance to directly judge the matching target. 2) Through the synchronous dynamic update of the template pixels themselves and their statistical characteristics (mean, variance), the template can actively track and adapt to the gradual changes of the target (such as the slow dust accumulation and color decline on the surface of the photovoltaic panel due to long-term use), and at the same time can effectively resist the mis-update caused by transient interference factors such as short-term occlusion and sudden light changes, thus ensuring the long-term effectiveness and robustness of the template library.

[0137] Subsequently, for emergency risk judgment and preliminary diversion at the edge end, if the condition of "the matching scores of at least one template and the spatial composition contour map reach the preset first threshold for at least two consecutive frames" is not met, it is determined that there is a risk for the current photovoltaic module. The edge device immediately triggers the emergency stop mechanism of the cleaning robot to prevent structural risks from causing damage to the robot or accidents. Report the risk status to the monitoring platform and terminate the current detection process; if the matching scores of at least one template and the spatial composition contour map reach the preset first threshold for at least two consecutive frames, it is determined that there is no obvious missing risk in the current detection area, and mark the preprocessed image as to be further analyzed. Subsequently, the system will execute the subsequent S5 cloud refined defect detection or S6 edge-end standby detection mechanism according to the network connection status between the current edge device and the cloud server. The cleaning robot can continue to move along the predetermined path to collect images of the next photovoltaic module.

[0138] Immediately afterwards, for the image confirmed to have no missing situation, perform the second link of perspective transformation image preprocessing (such as Figure 5 - IPM transformation): For the perspective distortion that may be brought by the wide-angle lens, using the pre-measured internal parameters of the camera and the external parameters of the camera relative to the photovoltaic panel plane, a fixed homography matrix H can be calculated ipm。Through this transformation matrix, a part of the ROI image after the first round of preprocessing (and passing the S3 integrity screening) is corrected from the original perspective view to a top view (bird's-eye view) without perspective effect. The key to this step is that it is carried out after the preliminary component integrity screening, which can avoid the problem of inaccurate extraction of feature points (such as corner points) used for real-time calculation of the transformation matrix due to the actual absence or severe deformation of the target photovoltaic panel components, thus avoiding the failure of preprocessing. Directly use the known internal and external camera parameters and the fixed transformation matrix H ipm for transformation without real-time corner detection, which improves the robustness and processing efficiency of image correction. Let the transformation matrix be H ipm . The inverse perspective formula is as follows:

[0139]

[0140] where (W std , H std ) are the normalized output image width and height, obtaining its "top view" on the photovoltaic panel plane. Select the pixel coordinates of the ROI in the original image, and then use the homography matrix for transformation without corner detection.

[0141] When the network connection status between the edge device and the cloud server is judged to be normal, VLM refined defect processing is carried out, and the edge device uploads the second-round preprocessed image (i.e., after inverse perspective correction) to the cloud server. There is a large visual language model (VLM) deployed in the cloud, such as a model of the Qwen-VL series (such as Qwen-VL-Max). To improve its performance in the photovoltaic defect detection task, this VLM model can be pre-instructed and fine-tuned using a dataset containing photovoltaic component images and their corresponding defect descriptions (text) to enhance the model's understanding of photovoltaic specific defect scenarios and cross-modal alignment capabilities.

[0142] After receiving the image to be detected, VLM will combine the preset and optimized photovoltaic defect semantic information (input through the dynamic prompt injection mechanism) to conduct detailed defect detection and classification on the image, mainly identifying two types of defects such as glass breakage and stains (bird droppings, dirt). These prompt words can include multiple dimensions, such as:

[0143] 1) System prompt (setting the role and task): "You are a professional photovoltaic module inspection assistant. You need to carefully analyze two types of defects in the photovoltaic module in the picture: (bird droppings, dirt) stains and glass breakage. You must answer strictly according to the following format: 1. First, describe in detail the phenomena you observe. 2. Then give the final judgment in the format of [X,X], where the first X represents the breakage condition (A = there is breakage, B = no breakage), and the second X represents the stain condition (A = there is a stain, B = no stain). If you are unsure about a certain judgment, use B. For example: [B,B] means no breakage and no stain."

[0144] 2) Initial question prompt (asking questions about specific defects): "Please carefully analyze this picture of the photovoltaic module and answer according to the following steps: 1. Describe in detail the phenomena you observe. 2. Ignore the background environment and give the final judgment on whether there are collapses and breakage stains on the photovoltaic panel in the format of [X,X]..."

[0145] 3) Spatial dimension prompt: Such as "Detect whether there are linear cracks in the upper left corner edge area of the photovoltaic module."

[0146] 4) Morphological dimension prompt: Such as "Find whether there are fragmented areas with a white spider web-like radiation morphology and blurred boundaries."

[0147] 5) Context / temporal dimension prompt (if there is historical data): Such as "Compare historical images to confirm whether the previously marked stain area has expanded." This dynamic prompt injection helps to guide the attention of the VLM to specific defect features, compensate for possible blind spots in edge detection, and improve the recognition accuracy of complex or subtle defects.

[0148] When analyzing, the VLM can selectively load cloud historical images or standard reference images related to the component to be detected. By comparing the similarity with the current input image, it helps to judge whether the defect is a newly added or developing defect, or helps to confirm the same type of defect pattern, further improving the accuracy of diagnosis. Finally, a standardized API interface is designed for data transmission (image upload, instruction issuance) and result parsing (defect type, location, confidence, etc.) between the edge device and the cloud VLM to ensure system integration.

[0149] When the network connection between the edge device and the cloud server is interrupted or unstable, the image information is saved in the local cache of the edge device, and a backup defect detection mechanism is executed to obtain a second defect detection result as a supplementary verification means for the cloud detection result to ensure that the basic defect detection does not interrupt.

[0150] For network instability or as a supplement to the cloud detection result, the following backup detection is performed on the edge device:

[0151] Stain defect detection: The input is the image of the photovoltaic module that has been subjected to inverse perspective correction, illumination compensation, and denoising. The RGB color image is converted into the HSV color space, and the pixel matrices of the H, S, and V channels are extracted respectively. The local threshold is calculated by the mean weighted method, and the window size for calculating the local threshold is defined to adapt to the non-uniform illumination on the surface of the photovoltaic panel. The multi-channel adaptive segmentation algorithm is used to distinguish the normal surface of the photovoltaic panel from the attachments such as stains and bird droppings based on the features of hue, saturation, and value. The three binary images are fused according to the logical rules, and the area jointly covered by the three is extracted according to the rules to improve the detection accuracy, or combined with logical OR according to the actual situation, which can ensure a relatively complete detection (such as Figure 4 ).

[0152] The defect detection results obtained by the edge-side backup detection mechanism will be temporarily stored in the local cache. When the network connection is restored, the previously cached images in the local will be successfully uploaded to the cloud VLM for more refined analysis to obtain the first defect detection result. After obtaining the detection result of the cloud VLM, the system will fuse the first defect detection result and the second defect detection result to obtain the final defect detection result.

[0153] The final defect detection results include two types of defects: stains and glass breakage; when fusing, first judge whether the first defect detection result is glass breakage; if so, the final defect detection result is glass breakage, otherwise, for the "stain" type of defect, the edge-side backup mechanism (the second defect detection result) is mainly based on the color space and is more sensitive to the "stain" type of defect (such as bird droppings, dirt). If the VLM detects a stain and the confidence level ≥ the third threshold, the VLM is adopted. If the VLM does not detect a stain or detects a stain but the confidence level < the third threshold, but the edge-side HSV color space analysis detects a stain signal, the result of the edge side is adopted. For the "glass breakage" type of defect, the color space backup mechanism on the edge side has a weak detection ability for "glass breakage", and such defects completely rely on the VLM.

[0154] To ensure the real-time performance of the system, an asynchronous pipeline architecture can be adopted during deployment. The tasks such as image acquisition, image preprocessing, and edge-side detection when the previous cleaning robot stops are executed in parallel with the image acquisition when the next stop occurs. A double-buffer queue or a similar mechanism can be used to manage the image data to be processed, thereby reducing the waiting time between tasks and achieving near-zero waiting data flow. If the edge device hardware supports (such as including an NPU), some edge-side model inference tasks (such as the computationally intensive parts of template matching or color segmentation) can be assigned to a dedicated hardware core for execution to further accelerate the processing. Similarly, the VLM inference on the cloud can also use a GPU or a dedicated AI accelerator to improve the efficiency.

[0155] The intelligent collaborative image defect detection mechanism described above significantly improves the stability and accuracy of defect detection in the complex and dynamic environments of photovoltaic power plants. This mechanism not only ensures accuracy for rapid industrial deployment but also effectively overcomes the shortcomings of existing single-detection technologies, which are susceptible to uncontrollable factors such as ambient lighting variations and insufficient dataset diversity. By leveraging the complementary advantages of multimodal and multi-stage approaches, it provides a highly robust and reliable intelligent photovoltaic module defect detection solution.

[0156] This invention effectively solves the problem of rapid and accurate identification of missing and deformed photovoltaic panels through an innovative dynamic template matching algorithm based on spatial position composition. Through a three-stage collaborative framework between the edge and the cloud, it cleverly combines the rapidity of template matching, the targeted nature of color space analysis, and the sophisticated analytical capabilities of visual language models. The dynamic template update mechanism enhances the system's robustness to gradual environmental changes and transient interference. The timing optimization of inverse perspective correction and specific optimization strategies for VLM (such as dynamic prompt injection) are both important innovative components of this patented solution, together forming an efficient, reliable, lightweight photovoltaic module defect detection solution that adapts to dynamic and complex scenarios. This technical solution has good engineering feasibility and scalability, adapts to actual industrial application needs, and promotes the improvement of the intelligent and automated level of photovoltaic power station operation and maintenance.

[0157] Table 1 shows the comparison of the accuracy of detecting all types of defects using the VLM model (without using reference images), VLM (using reference images), HSV multi-channel adaptive algorithm, and Deficiency-TM (spatial composition contour template matching) in the present invention.

[0158] surface

[0159]

[0160] As can be seen from the table above, the system, which combines the advantages of all models, can achieve good detection accuracy for all categories in both disconnected and connected conditions. This demonstrates that the system can effectively handle the complex safety inspection tasks involved in on-site cleaning of photovoltaic power plants, significantly improving detection accuracy and reliability.

Claims

1. A lightweight end-to-end defect detection method for a hanging-rail type photovoltaic cleaning robot, characterized in that, Including the following steps: Step 1. The photovoltaic cleaning robot runs linearly in a cyclic manner on a long-strip photovoltaic panel array composed of photovoltaic panels arranged in a straight line according to the following preset rules: stop, run a distance L1, stop, run a distance L2; The photovoltaic cleaning robot includes an edge computing device and an image acquisition device. Adjust the viewing angle of the image acquisition device so that the width of the photovoltaic panel in its field of view is not less than the width of the photovoltaic cleaning robot, and the length of the photovoltaic panel in the field of view is not less than the total length L of n photovoltaic panels arranged in a straight line, and L is not less than the sum of L1 and L2; When stopped, obtain the original image of the photovoltaic panel surface and the position of the photovoltaic cleaning robot relative to the photovoltaic panel array as the position corresponding to the original image; The original image includes a trapezoidal image of photovoltaic panels arranged in a straight line with a length of L; Step 2. After the first-round preprocessing of the original image, obtain the ROI region through the 4 vertices of the trapezoidal image region; extract the border line image and grid line image of the photovoltaic panel within the ROI region to form a spatial composition contour map corresponding to the original image; Step 3. On the edge computing device, perform template matching of the spatial composition contour map with the template at the corresponding position and the template at the position with a distance L1 from the corresponding position in the preset template library respectively to determine whether there is a risk; Step 4. If there is a risk, the photovoltaic cleaning robot stops advancing; otherwise, according to the network connection situation, use a visual language model to perform defect detection to obtain a first defect detection result, or use the edge computing device to execute a backup defect detection mechanism to obtain a second defect detection result.

2. The method according to claim 1, characterized in that, The determination of whether there is a risk includes the following steps: Calculate the matching degree score using the normalized correlation coefficient method and make a judgment based on the matching degree score; Match the spatial composition contour map with the template at the corresponding position and the template at the position with a distance L1 from the corresponding position in the preset template library respectively. If the matching score of at least one template with the spatial composition contour map reaches a preset first threshold, it means that the photovoltaic panel in the original image has no missing deformation, otherwise it means there is a risk; The risk refers to the risk that the photovoltaic cleaning robot cannot clean normally or falls from the photovoltaic panel due to missing or deformed photovoltaic panels; The template includes: a spatial composition contour template generated based on the trapezoidal images of several non-defective and non-deformed photovoltaic panels arranged in a straight line with a length of L.

3. The method according to claim 2, wherein The calculation process of the matching degree score includes the following steps: Calculate the normalized cross-correlation coefficient R(x, y) of the spatial composition contour map I and each template T according to the following formula ; ; ; Among them, (x, y): the reference coordinates of the starting point of the current match in the spatial composition contour map I, (i, j): the offset relative to the reference coordinates, and the value range is i ∈ (0, w−1), j ∈ (0, h−1), where w and h are the width and height dimensions of the template T respectively, and I(x, y) represents the pixel value of a point in the sub-block area with the same size as the template image T and with (x, y) as the upper left corner point in the spatial composition contour map I; μ I is the gray mean value of the current sub-block area, μ T is the gray mean value of the template image T; I(x+i, y+j) represents a pixel value of a point in a sub-block with the same size as T o in I, the upper left corner of which is located at (x, y) and the lower right corner is located at (x+w-1, y+h-1), and this sub-block will be compared with T o pixel by pixel to calculate the similarity; T(i, j) corresponds to I(x+i, y+j) one by one and represents the pixel value at the offset (i, j) in the template image; the value range of R(x, y) is [-1, 1], and the closer the value is to 1, the higher the matching degree.

4. The method according to claim 1, characterized in that, Step 3 further includes the following steps: When the matching score is greater than the first preset threshold, extract the matching region with the highest matching score, calculate the Mahalanobis distance D between this region and the corresponding template; if D is less than the preset second threshold, use the weighted moving average method to fuse and update the pixel values, mean and variance of the corresponding template; if D is not less than the preset second threshold, retain the old template, and this image is marked as having no missing but containing instantaneous interference.

5. The method according to claim 1, characterized in that A polarizer is provided on the image acquisition device; the first-round preprocessing includes light compensation and noise filtering; The image acquisition, the first-round preprocessing, and the standby defect detection mechanism at the time when the cleaning robot stopped last time are executed in parallel with the image acquisition at the next stop, and a double-buffer queue is used to manage the image data to be processed, thereby reducing the waiting time between tasks.

6. The method according to claim 1, wherein In step 4, the step of performing defect detection using the vision-language model according to the network connection situation to obtain a first defect detection result, or performing a standby defect detection mechanism using an edge computing device to obtain a second defect detection result specifically includes: When the network connection between the edge computing device and the cloud server is normal, the images determined to be risk-free are input into the vision-language model in the cloud, and defect detection is performed on the images in combination with the photovoltaic defect semantic information to obtain a first defect detection result; When the network connection is abnormal, the images determined to be risk-free are cached on the edge computing device, and a standby detection mechanism is executed. The standby detection mechanism includes stain defect detection based on the color space to obtain a second defect detection result; and after the network connection is restored to normal, the cached images are uploaded to the vision-language model in the cloud to obtain a first defect detection result, and the first defect detection result and the second defect detection result are fused to obtain a final defect detection result.

7. The method according to claim 6, wherein Before the step of inputting the images determined to be risk-free into the vision-language model in the cloud, the following steps are further included: Perform a second-round preprocessing on the images determined to be risk-free, and the second-round preprocessing includes inverse perspective transformation.

8. The method according to claim 6, wherein The step of fusing the first defect detection result and the second defect detection result includes the following steps: The final defect detection result includes two types of defects: stains and glass breakage; Judge whether the first defect detection result is glass breakage; if so, the final defect detection result is glass breakage, otherwise: Judge whether the VLM detects stains; if not, adopt the second defect detection result, Judge whether the confidence level of the stains detected by the VLM is less than a third threshold; if so, adopt the second defect detection result, if not, adopt the first defect detection result.

9. The method according to claim 1, wherein In step 1, the value of n is 3, 4 or 5; In step 2, extracting the frame line image and the grid line image of the photovoltaic panel within the ROI region includes: passing the image of the ROI region through an edge detection algorithm and a Hough line detection algorithm to obtain the frame line image and the grid line image of the photovoltaic panel within the ROI region.

10. The method according to claim 1, wherein The vision-language model is the Qwen-VL model, and it is fine-tuned using a dataset containing photovoltaic module images and their corresponding defect descriptions, The defects include glass breakage, bird droppings, and dirt; The prompt word framework of the vision-language model includes the following dimensions: Spatial dimension: the relative position description of the defect on the component; Morphological dimension: the crack direction or the stain diffusion morphological characteristics; Temporal dimension: comparative analysis with historical detection results.

Citation Information

Patent Citations

  • Gear and rack transmission type automatic line-changing photovoltaic panel monitoring device and method

    CN111463960A

  • Photovoltaic cell panel surface defect positioning method and device based on artificial intelligence

    CN112288730A

  • Deep learning-based photovoltaic module positioning and defect detection method in infrared image

    CN115082455A

  • Splicing method for solar panel images

    CN116152068A

  • Automatic operation and maintenance control method for intelligent photovoltaic cleaning robot

    CN118627796A