Photovoltaic array fault diagnosis and location method based on big data and machine learning

By constructing a spatiotemporal multi-scale photovoltaic array topology map and a learnable GaBP diagnostic localization model, the problem of unstable reliability in photovoltaic array fault diagnosis in existing technologies is solved, and fault localization with high accuracy and reliability is achieved.

CN122490370APending Publication Date: 2026-07-31FUZHOU SHENYUAN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FUZHOU SHENYUAN TECHNOLOGY CO LTD
Filing Date
2026-06-01
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

Existing photovoltaic array fault diagnosis methods do not adequately consider changes in photovoltaic operating conditions, acquisition noise, and multi-scale device access relationships, resulting in unstable fault location and location reliability. They also lack the ability to jointly model spatiotemporal multi-scale photovoltaic array topology maps and dynamically adjust fault propagation paths.

Method used

Using big data and machine learning methods, a spatiotemporal multi-scale photovoltaic array topology map is constructed, multi-scale Gaussian evidence vectors and node noise posterior parameters are generated, and fault diagnosis is performed through a learnable GaBP diagnostic localization model. The fault diagnosis and localization results are generated by combining the fault propagation path backtracking layer.

Benefits of technology

It improves the robustness and accuracy of photovoltaic array fault diagnosis, enhances the accuracy and reliability of fault location determination, and has the advantages of strong noise resistance and timely early warning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122490370A_ABST
    Figure CN122490370A_ABST
Patent Text Reader

Abstract

This invention discloses a photovoltaic array fault diagnosis and location method based on big data and machine learning, relating to the field of photovoltaic array fault diagnosis and location. The method includes: collecting and preprocessing photovoltaic power plant operation data and equipment connection relationships; identifying photovoltaic operating conditions and generating operating condition calibration data; constructing a spatiotemporal multi-scale photovoltaic array topology map; generating multi-scale Gaussian evidence vectors and node noise posterior parameters; inputting a learnable message propagation layer to generate spatiotemporal fault evidence representations of nodes; inputting a posterior uncertainty location output layer to generate fault location, location reliability, and initial warning level; and inputting a fault propagation path backtracking layer to generate photovoltaic array fault diagnosis and location results. This invention employs big data and a learnable GaBP model to achieve accurate photovoltaic array fault location, possessing advantages such as strong noise resistance, high reliability, and timely warning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of photovoltaic array fault diagnosis and location, and in particular to a photovoltaic array fault diagnosis and location method based on big data and machine learning. Background Technology

[0002] Currently, in the field of photovoltaic array fault diagnosis and location, existing technologies typically rely on current, voltage, power, and environmental condition data from photovoltaic power plant operation data to make threshold judgments or conventional machine learning classifications to identify abnormal states of strings, combiner box branches, or inverter input channels. However, such methods do not adequately consider changes in photovoltaic operating conditions, acquisition noise, and multi-scale equipment connection relationships, and are prone to misjudging irradiance fluctuations, temperature changes, or data noise as real faults, resulting in unstable fault location and location reliability.

[0003] At the same time, existing technologies mostly rely on a single point in time or a single device level for fault diagnosis. They lack joint modeling of spatiotemporal multi-scale photovoltaic array topology, continuous acquisition time markers, and propagation edges. This makes it difficult to express the propagation relationship of fault evidence at the string level, combiner box level, and inverter level, and also makes it difficult to correct the initial warning level based on the fault propagation path. As a result, the fault diagnosis and location results lack path interpretation and the ability to dynamically adjust the warning level. Summary of the Invention

[0004] One objective of this invention is to propose a photovoltaic array fault diagnosis and location method based on big data and machine learning. This invention uses big data and a learnable GaBP model to achieve accurate fault location of photovoltaic arrays, and has the advantages of strong noise resistance, high reliability and timely early warning.

[0005] The photovoltaic array fault diagnosis and location method based on big data and machine learning according to embodiments of the present invention includes: Collect photovoltaic power plant operation data and equipment connection relationships, preprocess the photovoltaic power plant operation data, and generate standardized photovoltaic operation data; Identify photovoltaic operating conditions based on standardized photovoltaic operation data, and generate operating condition calibration operation data based on photovoltaic operating conditions and standardized photovoltaic operation data; Construct a spatiotemporal multi-scale photovoltaic array topology based on device access relationships; The operating residuals in the spatiotemporal multi-scale photovoltaic array topology diagram are calculated based on the operating condition calibration data, and multi-scale Gaussian evidence vectors and node noise posterior parameters are generated based on the operating residuals. The multi-scale Gaussian evidence vector, node noise posterior parameters, and spatiotemporal multi-scale photovoltaic array topology graph are input into the learnable message propagation layer of the learnable GaBP diagnostic localization model of the photovoltaic array. The propagation edge is determined based on the spatiotemporal multi-scale photovoltaic array topology graph, and Gaussian messages corresponding to the propagation edge are generated based on the multi-scale Gaussian evidence vector. The message accuracy of the Gaussian message is corrected based on the node noise posterior parameters, and the corrected Gaussian message is jointly updated to generate a spatiotemporal fault evidence representation of the node. The spatiotemporal fault evidence of nodes is input into the posterior uncertainty location layer to generate fault location, location confidence and initial warning level; The node spatiotemporal fault evidence representation, fault location, location credibility, and initial warning level are input into the fault propagation path backtracking layer to generate a fault propagation path. Based on the fault propagation path, the initial warning level is corrected to generate the photovoltaic array fault diagnosis and location results.

[0006] Optionally, the photovoltaic power station operation data includes electrical operation fields, environmental condition fields, theoretical benchmark fields, and acquisition time identifiers; the equipment access relationships include string numbers, combiner box branch numbers, inverter input channel numbers, and connection attribution relationships; and the preprocessing includes time alignment, missing field completion, invalid value removal, unit unification, and numerical standardization.

[0007] Optionally, the generation of the operating condition calibration data includes: Read standardized photovoltaic operation data, divide the standardized photovoltaic operation data into continuous time segments based on the acquisition time identifier, perform segment statistical characterization processing on the standardized photovoltaic operation data in each continuous time segment, and generate operating condition identification features. The photovoltaic operating conditions are determined based on the operating condition identification characteristics, and operating condition calibration operation data is generated based on the photovoltaic operating conditions and standardized photovoltaic operation data.

[0008] Optionally, the generation of the spatiotemporal multi-scale photovoltaic array topology map includes: Read the acquisition time identifier from the device access relationship and standardized photovoltaic operation data; A multi-scale node set is generated based on the device access relationship, and spatial topology edges are established from string-level nodes to combiner box-level nodes and from combiner box-level nodes to inverter-level nodes under the same acquisition time identifier based on the connection affiliation relationship. Based on the acquisition time identifier, establish a time propagation edge between the same multi-scale nodes corresponding to adjacent acquisition time identifiers; Establish cross-scale propagation edges between string-level nodes and combiner box-level nodes, and between combiner box-level nodes and inverter-level nodes, based on connection affiliation relationships; By associating multi-scale node sets, spatial topological edges, temporal propagation edges, and cross-scale propagation edges according to the acquisition time identifier, a spatiotemporal multi-scale photovoltaic array topology map is generated.

[0009] Optionally, the generation of the multi-scale Gaussian evidence vector and the nodal noise posterior parameters includes: Read the calibration electrical operation field, theoretical benchmark field, and acquisition time identifier from the operating condition calibration operation data, and read the multi-scale node set, spatial topology edge, temporal propagation edge, and cross-scale propagation edge from the spatiotemporal multi-scale photovoltaic array topology graph; The calibration electrical operation fields are matched with the multi-scale node set according to the acquisition time identifier. Based on the calibration electrical operation fields and theoretical benchmark fields corresponding to string-level nodes, combiner box-level nodes and inverter-level nodes, the operation residuals of each topology node are generated. Gaussian evidence vectors for each topology node are generated based on the operational residuals of each topology node. The Gaussian evidence vectors of each topology node are then arranged to generate multi-scale Gaussian evidence vectors. Calculate the fluctuation amplitude of the operating residual of the same topology node under continuous acquisition time markers and the difference in operating residual between adjacent topology nodes, estimate the noise intensity and noise confidence coefficient of each topology node, and generate node noise posterior parameters.

[0010] Optionally, the generation of the node spatiotemporal fault evidence representation includes: The multi-scale Gaussian evidence vector, node noise posterior parameters, and spatiotemporal multi-scale photovoltaic array topology graph are input into the learnable message propagation layer of the photovoltaic array learnable GaBP diagnostic localization model. The learnable message propagation layer includes a propagation edge reading unit, an edge message input construction unit, a learnable Gaussian message generation unit, a noise accuracy correction unit, and a joint posterior update unit. The photovoltaic array learnable GaBP diagnostic localization model includes a learnable message propagation layer, a posterior uncertainty localization output layer, and a fault propagation path backtracking layer. In the learnable message propagation layer, the spatial topology edges, temporal propagation edges, and cross-scale propagation edges in the spatiotemporal multi-scale photovoltaic array topology graph are read by the propagation edge reading unit, and the spatial topology edges, temporal propagation edges, and cross-scale propagation edges are determined as the propagation edge set. The propagation edge set and multi-scale Gaussian evidence vector are input into the side message input construction unit. The corresponding Gaussian evidence components are read according to the start node and end node of each propagation edge, and the side message input vector is generated based on the start node Gaussian evidence component, the end node Gaussian evidence component and the propagation edge type. The learnable Gaussian message generation unit generates Gaussian messages based on the side message input vector; In the noise accuracy correction unit, the corrected message accuracy is generated based on the node noise posterior parameters, and the corrected Gaussian message is generated based on the message mean and the corrected message accuracy. The joint posterior update unit updates the termination node according to the Gaussian message after aggregation and correction of the termination node of the propagation edge, and generates a spatiotemporal fault evidence representation of the node.

[0011] Optionally, the generation of the fault location, location reliability, and initial warning level includes: The spatiotemporal fault evidence representation of the node is input into the posterior uncertainty localization output layer, which includes a candidate localization scoring unit, a fault location and credibility generation unit, and an initial warning level generation unit. In the posterior uncertainty positioning output layer, the spatiotemporal fault evidence representation of the node is read through the candidate positioning scoring unit to generate a candidate positioning node set, and a candidate position score is generated based on each candidate positioning node in the candidate positioning node set. The fault location and credibility generation unit determines the fault location by comparing the candidate location scores of each candidate location node in the candidate location node set, and generates location credibility based on the fault location. The fault location and location reliability are input into the initial warning level generation unit to generate the initial warning level.

[0012] Optionally, the generation of the photovoltaic array fault diagnosis and location results includes: The spatiotemporal fault evidence representation of the node, the fault location, the location credibility, and the initial warning level are input into the fault propagation path backtracking layer. The fault propagation path backtracking layer includes a path starting point determination unit, a path reverse expansion unit, a path propagation intensity calculation unit, and a warning level correction unit. In the fault propagation path backtracking layer, the path starting point determination unit reads the fault location and node spatiotemporal fault evidence representation, and determines the access object number and collection time identifier corresponding to the fault location as the path starting point. The spatiotemporal fault evidence of the node is input into the path backward expansion unit. Starting from the path origin, associated nodes are filtered and a candidate backtracking node sequence is generated. The path propagation intensity calculation unit generates path propagation intensity based on adjacent nodes in the candidate backtracking node sequence, and determines the candidate backtracking node sequence with the highest path propagation intensity as the fault propagation path; The warning level correction unit takes the initial warning level, location reliability, and path propagation strength as inputs, generates a warning level correction amount based on the location reliability and path propagation strength, corrects the initial warning level using the warning level correction amount, generates a corrected warning level, and combines the fault location, location reliability, fault propagation path, and corrected warning level into a photovoltaic array fault diagnosis and location result.

[0013] The beneficial effects of this invention are: This invention proposes a photovoltaic array fault diagnosis and localization method based on big data and machine learning. It constructs a spatiotemporal multi-scale photovoltaic array topology map based on device access relationships and calculates operational residuals using calibration data. Furthermore, it generates multi-scale Gaussian evidence vectors and node noise posterior parameters, enabling fault diagnosis to move beyond single device or single time point limitations. The multi-scale Gaussian evidence vectors can express anomalies at the string, combiner box, and inverter levels, while the node noise posterior parameters characterize the noise impact of different node data, providing a reliable basis for subsequent Gaussian message propagation and thus improving the robustness of fault diagnosis under complex field data.

[0014] This invention also inputs multi-scale Gaussian evidence vectors, node noise posterior parameters, and a spatiotemporal multi-scale photovoltaic array topology map into the learnable message propagation layer of the photovoltaic array learnable GaBP diagnostic localization model. The spatiotemporal multi-scale photovoltaic array topology map is used to determine propagation edges, and the node noise posterior parameters are used to correct the message accuracy of the Gaussian messages. The corrected Gaussian messages are then jointly updated to generate a spatiotemporal fault evidence representation for the nodes. This process allows spatial topological relationships, temporal evolution relationships, multi-scale device relationships, and noise reliability to work synergistically during the same propagation process, improving the accuracy and reliability of fault location determination. Attached Figure Description

[0015] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is an overall flowchart of the photovoltaic array fault diagnosis and location method based on big data and machine learning proposed in this invention; Figure 2 This is a flowchart of the learning message propagation layer in the photovoltaic array learning GaBP diagnostic and localization model of the photovoltaic array fault diagnosis and localization method based on big data and machine learning proposed in this invention. Figure 3 This is a flowchart of the fault propagation path backtracking layer in the photovoltaic array learnable GaBP diagnostic and localization model of the photovoltaic array fault diagnosis and localization method based on big data and machine learning proposed in this invention. Detailed Implementation

[0016] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0017] refer to Figures 1-3A photovoltaic array fault diagnosis and location method based on big data and machine learning includes: Collect photovoltaic power plant operation data and equipment connection relationships, preprocess the photovoltaic power plant operation data, and generate standardized photovoltaic operation data; Identify photovoltaic operating conditions based on standardized photovoltaic operation data, and generate operating condition calibration operation data based on photovoltaic operating conditions and standardized photovoltaic operation data; Construct a spatiotemporal multi-scale photovoltaic array topology based on device access relationships; The operating residuals in the spatiotemporal multi-scale photovoltaic array topology diagram are calculated based on the operating condition calibration data, and multi-scale Gaussian evidence vectors and node noise posterior parameters are generated based on the operating residuals. The multi-scale Gaussian evidence vector, node noise posterior parameters, and spatiotemporal multi-scale photovoltaic array topology graph are input into the learnable message propagation layer of the learnable GaBP diagnostic localization model of the photovoltaic array. The propagation edge is determined based on the spatiotemporal multi-scale photovoltaic array topology graph, and Gaussian messages corresponding to the propagation edge are generated based on the multi-scale Gaussian evidence vector. The message accuracy of the Gaussian message is corrected based on the node noise posterior parameters, and the corrected Gaussian message is jointly updated to generate a spatiotemporal fault evidence representation of the node. The spatiotemporal fault evidence of nodes is input into the posterior uncertainty location layer to generate fault location, location confidence and initial warning level; The node spatiotemporal fault evidence representation, fault location, location credibility, and initial warning level are input into the fault propagation path backtracking layer to generate a fault propagation path. Based on the fault propagation path, the initial warning level is corrected to generate the photovoltaic array fault diagnosis and location results.

[0018] In this embodiment, the photovoltaic power station operation data includes electrical operation fields, environmental condition fields, theoretical benchmark fields, and acquisition time identifiers. Specifically, the electrical operation fields include string current data, string voltage data, combiner box branch power data, inverter input power data, and inverter output power data. The environmental condition fields specifically include irradiance data, module temperature data, and ambient temperature data. The theoretical benchmark fields specifically include string benchmark current data, string benchmark voltage data, combiner box branch benchmark power data, and inverter input benchmark power data. The equipment access relationship includes string number, combiner box branch number, inverter input channel number, and connection affiliation. Preprocessing includes time alignment, missing field completion, invalid value removal, unit unification, and numerical standardization.

[0019] In this embodiment, the generation of operating condition calibration data includes: Read standardized photovoltaic operation data, divide the standardized photovoltaic operation data into continuous time segments based on the acquisition time identifier, perform segment statistical characterization processing on the standardized photovoltaic operation data in each continuous time segment, and generate operating condition identification features. The generation of operating condition identification features specifically includes: reading the standardized values ​​of irradiance, module temperature, ambient temperature, and inverter output power from standardized photovoltaic operation data; forming continuous time segments in order of acquisition time markers; averaging the standardized values ​​of irradiance within each continuous time segment to obtain the average irradiance; averaging the standardized values ​​of module temperature to obtain the average module temperature; averaging the standardized values ​​of ambient temperature to obtain the average ambient temperature; subtracting the standardized values ​​of irradiance under adjacent acquisition time markers to obtain the irradiance change; subtracting the standardized values ​​of module temperature under adjacent acquisition time markers to obtain the module temperature change; and arranging the average irradiance, average module temperature, average ambient temperature, irradiance change, and module temperature change in sequence into a one-dimensional vector, which is the operating condition identification feature. The photovoltaic operating condition is determined based on the operating condition identification characteristics, and operating condition calibration operation data is generated based on the photovoltaic operating condition and standardized photovoltaic operation data. The determination of photovoltaic operating conditions specifically includes: performing weighted summation and superimposed bias on the values ​​of each field in the operating condition identification features to generate operating condition scores for each candidate photovoltaic operating condition; performing exponential normalization on the operating condition scores for each candidate photovoltaic operating condition to generate operating condition probabilities for each candidate photovoltaic operating condition; and determining the candidate photovoltaic operating condition corresponding to the maximum operating condition probability as the photovoltaic operating condition. The generation of operating condition calibration data specifically includes: reading the standardized values ​​of string current, string voltage, combiner box branch power, and inverter input power from the standardized photovoltaic operating data; multiplying the standardized string current value by a current scaling factor and adding a current shift factor to generate the calibration string current value; multiplying the standardized string voltage value by a voltage scaling factor and adding a voltage shift factor to generate the calibration string voltage value; multiplying the standardized combiner box branch power value and the standardized inverter input power value by a power scaling factor and adding a power shift factor to generate the calibration combiner box branch power value and the calibration inverter input power value; establishing data rows according to the access object number and the acquisition time identifier, where the access object number specifically includes the string number, combiner box branch number, and inverter input channel number; and writing the calibration string current value, calibration string voltage value, calibration combiner box branch power value, and calibration inverter input power value into the corresponding field positions of the data row to generate the operating condition calibration data. The specific steps of calling the operating condition calibration coefficients include: calling the operating condition calibration coefficient table, which uses the photovoltaic operating condition as the index field and the electrical quantity scaling coefficient and electrical quantity translation coefficient as the value fields; performing index matching in the operating condition calibration coefficient table according to the photovoltaic operating condition; and reading the corresponding electrical quantity scaling coefficient and electrical quantity translation coefficient. The electrical quantity scaling coefficient includes the current scaling coefficient, voltage scaling coefficient, and power scaling coefficient, and the electrical quantity translation coefficient includes the current translation coefficient, voltage translation coefficient, and power translation coefficient. The generation of the operating condition calibration coefficient table specifically includes: reading the historical photovoltaic operating conditions, historical electrical operating fields, and historical theoretical benchmark fields from the historical normal operation data; grouping the historical electrical operating fields and historical theoretical benchmark fields according to the historical photovoltaic operating conditions; under each historical photovoltaic operating condition, using the historical electrical operating fields as input and the historical theoretical benchmark fields as output, performing least squares error fitting to obtain the electrical quantity scaling coefficient and electrical quantity translation coefficient corresponding to that historical photovoltaic operating condition; writing the historical photovoltaic operating conditions, electrical quantity scaling coefficient, and electrical quantity translation coefficient into the same data row to generate the operating condition calibration coefficient table.

[0020] In this embodiment, the generation of the spatiotemporal multi-scale photovoltaic array topology map includes: Read the acquisition time identifier from the device access relationship and standardized photovoltaic operation data; A multi-scale node set is generated based on the device access relationship, and spatial topology edges are established from string-level nodes to combiner box-level nodes and from combiner box-level nodes to inverter-level nodes under the same acquisition time identifier based on the connection affiliation relationship. The generation of the multi-scale node set specifically includes: reading the string number, combiner box branch number, and inverter input channel number from the device access relationship; mapping each string number to a string-level node; mapping each combiner box branch number to a combiner box-level node; and mapping each inverter input channel number to an inverter-level node. The string-level nodes, combiner box-level nodes, and inverter-level nodes are also called topology nodes. The nodes mapped from the string numbers are written to the string-level node scale identifier, and the corresponding string number is written to the access object number of that node. Similarly, the nodes mapped from the combiner box branch numbers are written to the combiner box-level node scale identifier, and the corresponding combiner box branch number is written to the access object number of that node. Finally, the nodes mapped from the inverter input channel numbers are written to the inverter-level node scale identifier, and the corresponding inverter input channel number is written to the access object number of that node, thus generating the multi-scale node set. The generation of spatial topology edges specifically includes: reading the attribution record between the string number and the combiner box branch number in the connection attribution relationship, pointing the corresponding string-level node to the corresponding combiner box-level node, generating a spatial topology edge from the string to the combiner box; reading the attribution record between the combiner box branch number and the inverter input channel number in the connection attribution relationship, pointing the corresponding combiner box-level node to the corresponding inverter-level node, generating a spatial topology edge from the combiner box to the inverter; and writing the acquisition time identifier into the spatial topology edge so that the spatial topology edge corresponds to the operating condition calibration data under the same acquisition time identifier. Based on the acquisition time identifier, establish a time propagation edge between the same multi-scale nodes corresponding to adjacent acquisition time identifiers; The generation of time propagation edges specifically includes: reading the acquisition time identifiers in the standardized photovoltaic operation data, sorting the acquisition time identifiers in chronological order to obtain a time identifier sequence, reading each node in the multi-scale node set, searching for two adjacent acquisition time identifiers in the time identifier sequence according to the access object number of the node, taking the node under the previous acquisition time identifier as the starting node of the time edge, taking the same node under the next acquisition time identifier as the ending node of the time edge, and writing the edge type as a time propagation edge. The above processing is performed on string-level nodes, combiner box-level nodes, and inverter-level nodes in the multi-scale node set respectively to generate time propagation edges. Establish cross-scale propagation edges between string-level nodes and combiner box-level nodes, and between combiner box-level nodes and inverter-level nodes, based on connection affiliation relationships; The generation of cross-scale propagation edges specifically includes: reading the corresponding records of string numbers and combiner box branch numbers in the connection attribution relationship; finding string-level nodes whose access object numbers are equal to the string number; finding combiner box-level nodes whose access object numbers are equal to the combiner box branch numbers; using the string-level nodes as the starting nodes and the combiner box-level nodes as the ending nodes to generate uplink cross-scale propagation edges; and using the combiner box-level nodes as the starting nodes and the string-level nodes as the ending nodes to generate downlink cross-scale propagation edges. This involves reading the combiner box branch numbers and inverter input channel numbers from the connection attribution relationship. Find the corresponding record of the access object number, find the combiner box level node whose access object number is equal to the combiner box branch number, and find the inverter level node whose access object number is equal to the inverter input channel number. Use the combiner box level node as the starting node and the inverter level node as the ending node to generate an uplink cross-scale propagation edge. Use the inverter level node as the starting node and the combiner box level node as the ending node to generate a downlink cross-scale propagation edge. Define the edges between the string level node and the combiner box level node, and between the combiner box level node and the inverter level node, used for evidence transmission at different scales as cross-scale propagation edges. The multi-scale node set, spatial topology edge, temporal propagation edge, and cross-scale propagation edge are associated according to the acquisition time identifier to generate a spatiotemporal multi-scale photovoltaic array topology map; The generation of the spatiotemporal multi-scale photovoltaic array topology graph specifically includes: reading the multi-scale node set, spatial topology edges, temporal propagation edges, and cross-scale propagation edges; writing node identifier, node scale identifier, access object number, and acquisition time identifier for each node in the multi-scale node set; writing start node identifier, end node identifier, edge type, and acquisition time identifier for each edge in the spatial topology edge, temporal propagation edge, and cross-scale propagation edge; merging the written multi-scale node set and the three types of edges into a graph structure data table to generate the spatiotemporal multi-scale photovoltaic array topology graph.

[0021] In this embodiment, the generation of the multi-scale Gaussian evidence vector and the nodal noise posterior parameters includes: Read the calibration electrical operation field, theoretical benchmark field, and acquisition time identifier from the operating condition calibration operation data, and read the multi-scale node set, spatial topology edge, temporal propagation edge, and cross-scale propagation edge from the spatiotemporal multi-scale photovoltaic array topology graph; The calibration electrical operation fields are matched with the multi-scale node set according to the acquisition time identifier. Based on the calibration electrical operation fields and theoretical benchmark fields corresponding to string-level nodes, combiner box-level nodes and inverter-level nodes, the operation residuals of each topology node are generated. The generation of operational residuals specifically includes: reading the calibration electrical operation field and theoretical reference field from the operating condition calibration data. The calibration electrical operation field includes the calibration string current value, calibration string voltage value, calibration combiner box branch power value, and calibration inverter input power value. The theoretical reference field includes string reference current data, string reference voltage data, combiner box branch reference power data, and inverter input reference power data. The calibration electrical operation field and theoretical reference field are paired according to the acquisition time identifier, and the string reference current data is subtracted from the calibration string current value. The string current operating residual is generated by dividing the string current value by the absolute value of the string reference current data. The string voltage operating residual is generated by subtracting the string reference voltage data from the calibrated string voltage value and then dividing it by the absolute value of the string reference voltage data. The combiner box branch power operating residual is generated by subtracting the combiner box branch reference power data from the calibrated combiner box branch power value and then dividing it by the absolute value of the combiner box branch reference power data. The inverter input power operating residual is generated by subtracting the inverter input reference power data from the calibrated inverter input power value and then dividing it by the absolute value of the inverter input reference power data. Gaussian evidence vectors for each topology node are generated based on the operational residuals of each topology node. The Gaussian evidence vectors of each topology node are then arranged to generate multi-scale Gaussian evidence vectors. The generation of multi-scale Gaussian evidence vectors specifically includes: reading the string current operation residual and string voltage operation residual corresponding to the string-level node, reading the combiner box branch power operation residual corresponding to the combiner box-level node, reading the inverter input power operation residual corresponding to the inverter-level node, performing standardized Gaussian parameter mapping on the operation residuals of each scale respectively, converting the operation residuals of each scale into standardized residual values, residual mean, residual variance and initial anomaly probability, arranging the standardized residual values, residual mean, residual variance and initial anomaly probability in a fixed field order to generate isomorphic scale Gaussian evidence components, organizing each isomorphic scale Gaussian evidence component according to the node scale identifier, access object number and acquisition time identifier to generate multi-scale Gaussian evidence vectors; The process of standardizing Gaussian parameter mapping specifically includes: reading the running residual of the topology node under the current acquisition time identifier; reading the historical residual mean and historical residual variance of the topology node under the corresponding photovoltaic operating condition; subtracting the running residual from the historical residual mean and scaling it out using the historical residual standard deviation corresponding to the historical residual variance to generate a standardized residual value; writing the historical residual mean into the residual mean field and the historical residual variance into the residual variance field; taking the absolute value of the standardized residual value and inputting the absolute value into the monotonic probability mapping function to generate the initial anomaly probability. The processing of the monotonic probability mapping function specifically includes: inputting the absolute value of the standardized residual into the exponential decay mapping to first generate the normal probability, and then subtracting the normal probability from 1 to obtain the initial abnormal probability; Calculate the fluctuation amplitude of the running residual of the same topology node under continuous acquisition time markers and the difference in running residual between adjacent topology nodes, estimate the noise intensity and noise confidence coefficient of each topology node, and generate node noise posterior parameters. The generation of node noise posterior parameters specifically includes: reading the operational residuals of the same topological node under continuous acquisition time markers, calculating the absolute value of the difference between the operational residuals under adjacent acquisition time markers, averaging the absolute values ​​of the differences to generate temporal noise intensity; reading the operational residuals of adjacent topological nodes connected by spatial topological edges and cross-scale propagation edges, calculating the absolute value of the difference between the operational residuals of adjacent topological nodes, averaging the absolute values ​​of the differences to generate structural noise intensity; normalizing and weighting the temporal noise intensity and structural noise intensity to generate noise intensity; performing inverse normalization on the noise intensity to generate noise confidence coefficient; and associating the noise intensity and noise confidence coefficient according to the topological node and acquisition time marker to generate node noise posterior parameters.

[0022] In this embodiment, the generation of node spatiotemporal fault evidence representation includes: The multi-scale Gaussian evidence vector, node noise posterior parameters, and spatiotemporal multi-scale photovoltaic array topology graph are input into the learnable message propagation layer of the photovoltaic array learnable GaBP diagnostic localization model. The learnable message propagation layer includes a propagation edge reading unit, an edge message input construction unit, a learnable Gaussian message generation unit, a noise accuracy correction unit, and a joint posterior update unit. The photovoltaic array learnable GaBP diagnostic localization model includes a learnable message propagation layer, a posterior uncertainty localization output layer, and a fault propagation path backtracking layer. The training process of the learnable GaBP diagnostic localization model for photovoltaic arrays includes: constructing a training sample set, which includes sample photovoltaic power plant operation data, sample equipment connection relationships, sample fault location labels, sample localization confidence labels, and sample early warning level labels; preprocessing the sample photovoltaic power plant operation data to generate standardized photovoltaic operation data; identifying sample photovoltaic operating conditions based on the standardized photovoltaic operation data, and generating sample operating condition calibration operation data based on the sample photovoltaic operating conditions and the standardized photovoltaic operation data; constructing a sample spatiotemporal multi-scale photovoltaic array topology map based on the sample equipment connection relationships; calculating the sample operation residuals in the sample spatiotemporal multi-scale photovoltaic array topology map based on the sample operating condition calibration operation data, and generating sample multi-scale Gaussian evidence vectors and sample node noise posterior parameters based on the sample operation residuals; and inputting the sample multi-scale Gaussian evidence vectors, sample node noise posterior parameters, and sample spatiotemporal multi-scale photovoltaic array topology map into the learnable GaBP model for photovoltaic arrays. The learnable message propagation layer of the diagnostic localization model generates spatiotemporal fault evidence representations of sample nodes. These representations are then input into the posterior uncertainty localization output layer to generate sample fault locations, sample localization confidence levels, and sample initial warning levels. Finally, these representations are input into the fault propagation path backtracking layer to generate sample fault propagation paths and sample photovoltaic array fault diagnosis and localization results. A training loss is constructed based on the differences between sample fault locations and their labels, between sample localization confidence levels and their labels, and between sample photovoltaic array fault diagnosis and localization results and their warning level labels. Backpropagation is used to update the learnable parameters in the learnable message propagation layer, the posterior uncertainty localization output layer, and the fault propagation path backtracking layer until the training loss meets the convergence condition, resulting in a trained learnable GaBP diagnostic localization model for the photovoltaic array. The improvements to this model include: First, the fixed Gaussian message propagation method in traditional GaBP is improved to a learnable message propagation method. Propagation edges are determined based on the spatiotemporal multi-scale photovoltaic array topology graph, Gaussian messages corresponding to the propagation edges are generated based on multi-scale Gaussian evidence vectors, and the message accuracy of Gaussian messages is corrected using node noise posterior parameters, thus reducing the impact of messages generated by high-noise nodes in joint updates. Second, the traditional single-scale topology inference is improved to spatiotemporal multi-scale joint inference, utilizing spatial topology edges, temporal propagation edges, and cross-scale propagation edges to participate in Gaussian message updates, enabling the model to simultaneously possess string-level fine-grained positioning, combiner-level mesoscale aggregation, inverter-level global consistency judgment, and cross-time fault evolution expression capabilities. Third, the node posterior inference results of traditional GaBP are extended to diagnostic positioning outputs for operation and maintenance early warning. Fault location, positioning credibility, and initial warning level are generated based on node spatiotemporal fault evidence representation, and fault propagation path is generated through a fault propagation path backtracking layer. The initial warning level is then corrected based on the fault propagation path, expanding the model output from a single anomaly probability to a closed-loop result of fault location, credibility assessment, propagation path interpretation, and warning level generation. In the learnable message propagation layer, the spatial topology edges, temporal propagation edges, and cross-scale propagation edges in the spatiotemporal multi-scale photovoltaic array topology graph are read by the propagation edge reading unit, and the spatial topology edges, temporal propagation edges, and cross-scale propagation edges are determined as the propagation edge set. The generation of the propagation edge set specifically includes: reading the edge type field, start node field, end node field, and acquisition time identifier field from the spatiotemporal multi-scale photovoltaic array topology graph; writing the records of spatial topology edges corresponding to the edge type field into the spatial propagation subset; writing the records of temporal propagation edges corresponding to the edge type field into the temporal propagation subset; writing the records of cross-scale propagation edges corresponding to the edge type field into the cross-scale propagation subset; and merging the spatial propagation subset, temporal propagation subset, and cross-scale propagation subset to generate the propagation edge set. The propagation edge set and multi-scale Gaussian evidence vector are input into the side message input construction unit. The corresponding Gaussian evidence components are read according to the start node and end node of each propagation edge, and the side message input vector is generated based on the start node Gaussian evidence component, the end node Gaussian evidence component and the propagation edge type. The generation of the side message input vector specifically includes: reading each propagation edge in the propagation edge set; reading the starting node Gaussian evidence component from the multi-scale Gaussian evidence vector according to the starting node field of the propagation edge; reading the ending node Gaussian evidence component from the multi-scale Gaussian evidence vector according to the ending node field of the propagation edge; subtracting the standardized residual value in the starting node Gaussian evidence component from the standardized residual value in the ending node Gaussian evidence component to generate the node residual difference value; converting the propagation edge type into the edge type code; and arranging the node residual difference value, the initial anomaly probability of the starting node, the initial anomaly probability of the ending node, and the edge type code in order to generate the side message input vector. The learnable Gaussian message generation unit generates Gaussian messages based on the side message input vector; The generation of Gaussian messages specifically includes: reading the edge message input vector, performing weighted summation and superposition bias on the field values ​​in the edge message input vector to generate the message intermediate vector, performing nonlinear activation processing on the message intermediate vector, mapping the message intermediate vector after nonlinear activation processing to the message mean and the original message precision, and generating a Gaussian message. A Gaussian message is a binary message vector group composed of the message mean and the original message precision. The generation of the message mean specifically includes: inputting the message intermediate vector after nonlinear activation processing into the mean projection branch; the mean projection branch performs a weighted summation on the message intermediate vector and superimposes the mean bias to generate the message mean, where the message mean is used to represent the location of the fault evidence center transmitted to the termination node by a certain propagation edge; when the edge message input vector corresponds to only one node evidence component, the message mean is in numerical form; when the edge message input vector contains multiple evidence fields at the same time, the message mean is in vector form. The generation of the original message precision specifically includes: inputting the message intermediate vector after nonlinear activation processing into the precision projection branch, performing a weighted summation on the message intermediate vector and superimposing the precision bias, and then performing positive value processing to generate the original message precision, where the original message precision is used to represent the reliability of the Gaussian message, and is equal to the reciprocal of the message variance. In the noise accuracy correction unit, the corrected message accuracy is generated based on the node noise posterior parameters, and the corrected Gaussian message is generated based on the message mean and the corrected message accuracy. The generation of the corrected Gaussian message specifically includes: reading the noise confidence coefficient in the posterior parameters of the node noise, multiplying the noise confidence coefficient corresponding to the starting node of the propagation edge with the original message precision in the Gaussian message to generate the corrected message precision, retaining the message mean in the Gaussian message, and arranging the message mean and the corrected message precision in order to generate the corrected Gaussian message. The joint posterior update unit updates the termination node according to the Gaussian message after aggregation and correction of the termination node of the propagation edge, and generates a spatiotemporal fault evidence representation of the node. The generation of node spatiotemporal fault evidence representation specifically includes: grouping the corrected Gaussian messages according to the terminating node; calculating the weighted average of the corrected Gaussian messages received by the same terminating node according to the spatial topology edge, temporal propagation edge, and cross-scale propagation edge to generate spatial evidence items, temporal evidence items, and cross-scale evidence items; summing and fusing the spatial evidence items, temporal evidence items, and cross-scale evidence items, and performing residual connection with the original Gaussian evidence components of the terminating node to generate a fused posterior vector; mapping the fused posterior vector to obtain posterior mean candidate values, posterior variance candidate values, and fault score values, where the posterior mean candidate values, posterior variance candidate values, and fault score values ​​are all obtained by performing weighted summation and superimposing bias on the fused posterior vector; using the posterior mean candidate values ​​as the posterior mean; performing positive value processing on the posterior variance candidate values ​​to generate posterior variance; performing probability normalization processing on the fault score values ​​to generate fault confidence; arranging the posterior mean, posterior variance, and fault confidence values ​​according to the access object number and collection time identifier to generate node spatiotemporal fault evidence representation.

[0023] In this embodiment, the generation of fault location, location reliability, and initial warning level includes: The spatiotemporal fault evidence representation of the node is input into the posterior uncertainty localization output layer, which includes a candidate localization scoring unit, a fault location and credibility generation unit, and an initial warning level generation unit. In the posterior uncertainty positioning output layer, the spatiotemporal fault evidence representation of the node is read through the candidate positioning scoring unit to generate a candidate positioning node set, and a candidate position score is generated based on each candidate positioning node in the candidate positioning node set. The generation of the candidate location node set specifically includes: reading the posterior mean, posterior variance, fault confidence, access object number, node scale identifier, and collection time identifier from the spatiotemporal fault evidence representation of the node; establishing candidate node records according to the same access object number and the same collection time identifier; writing the posterior mean, posterior variance, fault confidence, and node scale identifier into the corresponding candidate node record; and generating the candidate location node set. The generation of candidate location scores specifically includes: reading the posterior mean, posterior variance, and fault confidence of each candidate location node in the candidate location node set; taking the absolute value of the posterior mean and normalizing it to generate a normalized absolute value of the posterior mean; taking the reciprocal of the posterior variance and normalizing it to generate a variance confidence value; normalizing the fault confidence to generate a confidence normalized value; and multiplying the confidence normalized value, the normalized absolute value of the posterior mean, and the variance confidence value to generate a candidate location score. The fault location and credibility generation unit determines the fault location by comparing the candidate location scores of each candidate location node in the candidate location node set, and generates location credibility based on the fault location. The generation of the fault location specifically includes: reading the candidate location score of each candidate location node in the candidate location node set, sorting the candidate location scores from largest to smallest, and determining the access object number and node scale identifier corresponding to the candidate location node with the highest score as the fault location; The generation of location confidence specifically includes: reading the fault confidence and posterior variance corresponding to the fault location, taking the reciprocal of the posterior variance and normalizing it to generate the fault location variance confidence value, and multiplying the fault confidence corresponding to the fault location with the fault location variance confidence value to generate the location confidence. The fault location and location reliability are input into the initial warning level generation unit to generate the initial warning level; The generation of the initial warning level specifically includes: reading the node scale identifier and location reliability corresponding to the fault location; when the node scale identifier is a string-level node scale identifier, generating a string-level initial warning level according to the location reliability; when the node scale identifier is a combiner box-level node scale identifier, generating a combiner box-level initial warning level according to the location reliability; when the node scale identifier is an inverter-level node scale identifier, generating an inverter-level initial warning level according to the location reliability; and using the string-level initial warning level, combiner box-level initial warning level, and inverter-level initial warning level as the initial warning level. The generation of the initial warning level specifically includes the following steps: The initial warning level is determined using a five-stage method. When the location reliability is greater than or equal to 0 and less than 0.2, a Level 1 candidate warning level is generated; when the location reliability is greater than or equal to 0.2 and less than 0.4, a Level 2 candidate warning level is generated; when the location reliability is greater than or equal to 0.4 and less than 0.6, a Level 3 candidate warning level is generated; when the location reliability is greater than or equal to 0.6 and less than 0.8, a Level 4 candidate warning level is generated; and when the location reliability is greater than or equal to 0.8 and less than or equal to 1, a Level 5 candidate warning level is generated. Then, a level increase process is performed based on the node scale identifier. If the increased level is greater than the Level 5 warning level, the Level 5 warning level is used as the initial warning level. Specifically, when the node scale identifier is a string-level node scale identifier, the candidate warning level is used as the initial warning level; when the node scale identifier is a combiner box-level node scale identifier, the candidate warning level is increased by one level and used as the initial warning level; and when the node scale identifier is an inverter-level node scale identifier, the candidate warning level is increased by two levels and used as the initial warning level.

[0024] In this embodiment, the generation of photovoltaic array fault diagnosis and location results includes: The spatiotemporal fault evidence representation of the node, the fault location, the location credibility, and the initial warning level are input into the fault propagation path backtracking layer. The fault propagation path backtracking layer includes a path starting point determination unit, a path reverse expansion unit, a path propagation intensity calculation unit, and a warning level correction unit. In the fault propagation path backtracking layer, the path starting point determination unit reads the fault location and node spatiotemporal fault evidence representation, and determines the access object number and collection time identifier corresponding to the fault location as the path starting point. The generation of the path starting point specifically includes: reading the fault location, which includes the access object number and node scale identifier; reading the collection time identifier in the spatiotemporal fault evidence representation of the node; writing the access object number, node scale identifier and latest collection time identifier corresponding to the fault location into the same path starting point record; generating the path starting point; the path starting point is used by the path reverse expansion unit to generate a candidate backtracking node sequence. The spatiotemporal fault evidence of the node is input into the path backward expansion unit. Starting from the path origin, associated nodes are filtered and a candidate backtracking node sequence is generated. The generation of the candidate backtracking node sequence specifically includes: reading the path start point, the fault confidence in the spatiotemporal fault evidence representation of the node, the access object number, the node scale identifier, and the collection time identifier. Taking the path start point as the current node, the node with the collection time identifier earlier than the current node and the fault confidence is not greater than the fault confidence of the current node is searched. Among the found nodes, nodes with the same access object number, nodes with adjacent node scale identifiers, and nodes in the same collection time neighborhood as the current node are retained. The retained nodes are arranged from near to far according to the collection time identifier to generate the candidate backtracking node sequence. The path propagation intensity calculation unit generates path propagation intensity based on adjacent nodes in the candidate backtracking node sequence, and determines the candidate backtracking node sequence with the highest path propagation intensity as the fault propagation path; The generation of path propagation strength specifically includes: reading the fault confidence, posterior variance, and node scale identifier of two adjacent nodes in the candidate backtracking node sequence; subtracting the fault confidence of two adjacent nodes and taking the absolute value to generate the confidence propagation quantity; taking the reciprocal of the posterior variance of two adjacent nodes and averaging them to generate the variance reliable propagation quantity; generating the scale uplink propagation quantity when the node scale identifier of two adjacent nodes points from the string level to the combiner box level or from the combiner box level to the inverter level; generating the scale maintenance propagation quantity when the node scale identifier of two adjacent nodes remains consistent; weighted summing of the confidence propagation quantity, variance reliable propagation quantity, and scale propagation quantity to generate the adjacent node propagation strength; and averaging the propagation strengths of all adjacent nodes in the candidate backtracking node sequence to generate the path propagation strength. The generation of the fault propagation path specifically includes: generating the path propagation intensity for each candidate backtracking node sequence, comparing the path propagation intensity of each candidate backtracking node sequence, and determining the candidate backtracking node sequence corresponding to the maximum path propagation intensity as the fault propagation path. The fault propagation path includes the path starting point, backtracking nodes, the propagation intensity of each adjacent node, and the path propagation intensity. The warning level correction unit takes the initial warning level, location reliability, and path propagation strength as inputs, generates a warning level correction amount based on the location reliability and path propagation strength, corrects the initial warning level using the warning level correction amount, generates a corrected warning level, and combines the fault location, location reliability, fault propagation path, and corrected warning level into a photovoltaic array fault diagnosis and location result. The generation of the revised warning level specifically includes: reading the initial warning level, location reliability, and path propagation strength. Both location reliability and path propagation strength are normalized values ​​between 0 and 1. The location reliability is multiplied by the path propagation strength to generate a path reliability linkage value. When the path reliability linkage value is greater than or equal to 0 and less than 0.36, the warning level revision amount is assigned to 0. When the path reliability linkage value is greater than or equal to 0.36 and less than 0.64, the warning level revision amount is assigned to 1. When the path reliability linkage value is greater than or equal to 0.64 and less than or equal to 1, the warning level revision amount is assigned to 2. The initial warning level is added to the warning level revision amount to obtain the revised warning level. When the revised warning level is greater than the Level 5 warning level, the Level 5 warning level is used as the revised warning level. The generation of photovoltaic array fault diagnosis and location results specifically includes: writing the fault location, location reliability, fault propagation path, and corrected warning level into the same result record to generate photovoltaic array fault diagnosis and location results. The photovoltaic array fault diagnosis and location results are used to output the access object where the fault is located, the location reliability, the fault propagation path, and the corrected warning level.

[0025] Example 1: To verify the feasibility of this invention in practice, it was applied to an intelligent monitoring platform for a centralized photovoltaic power station in a coastal area. This power station is located in an area with frequent changes in sea breezes, salt spray, and cloud cover, and is equipped with a large array of photovoltaic panels, combiner boxes, inverters, and environmental monitoring devices. The platform continuously collects string current data, string voltage data, combiner box branch power data, inverter input power data, inverter output power data, irradiance data, module temperature data, ambient temperature data, and the time of collection. It also saves the string number, combiner box branch number, inverter input channel number, and connection relationship. A major long-standing problem in this scenario is that traditional threshold alarms are prone to false alarms when local cloud shadows pass, module temperatures change rapidly, or communication data fluctuates briefly. Furthermore, when a string, combiner box branch, or inverter input channel experiences a persistent anomaly, traditional platforms can only provide general alarms, requiring maintenance personnel to check curves and connection relationships step by step.

[0026] When applying this invention, the photovoltaic power plant operation data is first connected to the data processing link and preprocessed. Then, the photovoltaic operating conditions are identified based on the standardized photovoltaic operating data, and the changes in irradiance, module temperature, ambient temperature, and inverter output status are incorporated into the operating condition judgment. Based on the photovoltaic operating conditions, the standardized photovoltaic operating data is calibrated to generate operating condition calibration operating data.

[0027] After generating the operating condition calibration data, the system constructs a spatiotemporal multi-scale photovoltaic array topology map based on device access relationships. This topology map maps strings, combiner box branches, and inverter input channels to nodes of different scales, and establishes spatial topology edges based on connection affiliation, temporal propagation edges based on acquisition time identifiers, and cross-scale propagation edges based on the correspondence between different device levels. Next, the system calculates the operating residuals based on the operating condition calibration data and generates multi-scale Gaussian evidence vectors and node noise posterior parameters. The node noise posterior parameters describe the degree to which the current data of a node is affected by acquisition jitter, communication fluctuations, or environmental disturbances, thereby reducing the impact of high-noise nodes on diagnostic conclusions during subsequent propagation. Subsequently, the multi-scale Gaussian evidence vectors, node noise posterior parameters, and the spatiotemporal multi-scale photovoltaic array topology map are input into the learnable message propagation layer of the photovoltaic array learnable GaBP diagnostic localization model. This layer determines propagation edges based on the spatiotemporal multi-scale photovoltaic array topology map, generates Gaussian messages corresponding to the propagation edges based on the multi-scale Gaussian evidence vectors, and corrects the message accuracy of the Gaussian messages based on the node noise posterior parameters. The corrected Gaussian message is jointly updated on spatial topological edges, temporal propagation edges, and cross-scale propagation edges to generate a spatiotemporal fault evidence representation for the node. The posterior uncertainty localization output layer further generates the fault location, localization confidence level, and initial warning level, while the fault propagation path backtracking layer generates the fault propagation path and corrects the initial warning level based on the path propagation. Therefore, the photovoltaic array fault diagnosis and localization results output by the platform not only provide the access object where the fault occurs but also the localization confidence level, the fault propagation path, and the corrected warning level.

[0028] To facilitate evaluation of the implementation effect, a traditional fixed threshold alarm combined with manual topology inspection was used as a control method for comparison with the method of this invention. Data sources included records from the photovoltaic intelligent monitoring platform, inverter operation logs, combiner box branch data, environmental monitoring data, and operation and maintenance verification records. Comparison indicators included fault location accuracy, false alarm rate, missed alarm rate, average location time, and early warning lead time.

[0029] Table 1 Comparison of the Implementation Effects of Photovoltaic Array Fault Diagnosis and Location

[0030] As shown in Table 1, the method of this invention achieves a fault location accuracy of 96.4%, which is higher than the 82.6% of the traditional fixed threshold alarm combined with manual investigation method, and also higher than the 88.9% of the conventional machine learning classification method. This improvement mainly comes from the combined effect of operating condition calibration data and spatiotemporal multi-scale photovoltaic array topology map. Traditional methods usually use electrical quantity deviations from thresholds directly as alarm basis, making it difficult to distinguish between irradiance fluctuations, component temperature changes, and actual faults; although conventional machine learning classification methods can learn some abnormal features, they lack the propagation relationship between string level, combiner box level, and inverter level. This invention first generates operating residuals, and then converts the residuals into multi-scale Gaussian evidence vectors, so that abnormal evidence can be expressed in the equipment access relationship, thus making fault location judgment more stable.

[0031] Regarding the false alarm rate and false negative rate, the method of this invention has a false alarm rate of 3.6% and a false negative rate of 2.8%, respectively, which are significantly lower than the control method. The reason for the reduced false alarm rate is that the posterior parameters of node noise participate in the message accuracy correction of Gaussian messages. When the data of a certain node is affected by communication jitter or environmental disturbances, the impact of the Gaussian message generated by that node in the joint update is reduced, avoiding the amplification of single-point noise into fault alarms. The reduced false negative rate is related to the time propagation edge and the cross-scale propagation edge. The string anomalies that are not obvious in a short period of time can accumulate evidence along the continuous acquisition time markers and further propagate to the combiner box level and the inverter level, making early weak faults easier to identify.

[0032] Regarding the average location time, the traditional method requires 37.5 minutes, while the method of this invention reduces it to 8.6 minutes, demonstrating that this invention can significantly reduce the time spent on manual step-by-step troubleshooting. This is because the posterior uncertainty location output layer directly generates the fault location and location reliability, while the fault propagation path backtracking layer further provides the transmission relationship of anomaly evidence between nodes at different scales. Maintenance personnel can quickly determine the source of the anomaly based on the fault propagation path without repeatedly switching between string curves, combiner box branch records, and inverter logs.

[0033] In terms of early warning lead time and path interpretability coverage, the method of this invention achieves 32.4 minutes and 91.7% respectively, a significant improvement over traditional methods. This result demonstrates that the present invention can not only determine the current fault location but also express the fault evolution process using time propagation and the fault impact range using cross-scale propagation. After the initial warning level is corrected by the fault propagation path, it can better align with the fault source level, propagation intensity, and location reliability, transforming the warning result from a single-point alarm into an operational decision result with path-based information.

[0034] The above are merely preferred embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A method for photovoltaic array fault diagnosis and location based on big data and machine learning, characterized in that, include: Collect photovoltaic power plant operation data and equipment connection relationships, preprocess the photovoltaic power plant operation data, and generate standardized photovoltaic operation data; Identify photovoltaic operating conditions based on standardized photovoltaic operation data, and generate operating condition calibration operation data based on photovoltaic operating conditions and standardized photovoltaic operation data; Construct a spatiotemporal multi-scale photovoltaic array topology based on device access relationships; The operating residuals in the spatiotemporal multi-scale photovoltaic array topology diagram are calculated based on the operating condition calibration data, and multi-scale Gaussian evidence vectors and node noise posterior parameters are generated based on the operating residuals. The multi-scale Gaussian evidence vector, node noise posterior parameters, and spatiotemporal multi-scale photovoltaic array topology graph are input into the learnable message propagation layer of the learnable GaBP diagnostic localization model of the photovoltaic array. The propagation edge is determined based on the spatiotemporal multi-scale photovoltaic array topology graph, and Gaussian messages corresponding to the propagation edge are generated based on the multi-scale Gaussian evidence vector. The message accuracy of the Gaussian message is corrected based on the node noise posterior parameters, and the corrected Gaussian message is jointly updated to generate a spatiotemporal fault evidence representation of the node. The spatiotemporal fault evidence of nodes is input into the posterior uncertainty location layer to generate fault location, location confidence and initial warning level; The node spatiotemporal fault evidence representation, fault location, location credibility, and initial warning level are input into the fault propagation path backtracking layer to generate a fault propagation path. Based on the fault propagation path, the initial warning level is corrected to generate the photovoltaic array fault diagnosis and location results.

2. The photovoltaic array fault diagnosis and location method based on big data and machine learning according to claim 1, characterized in that, The photovoltaic power station operation data includes electrical operation fields, environmental condition fields, theoretical benchmark fields, and acquisition time identifiers. The equipment access relationships include string numbers, combiner box branch numbers, inverter input channel numbers, and connection attribution relationships. The preprocessing includes time alignment, missing field completion, invalid value removal, unit unification, and numerical standardization.

3. The photovoltaic array fault diagnosis and location method based on big data and machine learning according to claim 1, characterized in that, The generation of the operating condition calibration data includes: Read standardized photovoltaic operation data, divide the standardized photovoltaic operation data into continuous time segments based on the acquisition time identifier, perform segment statistical characterization processing on the standardized photovoltaic operation data in each continuous time segment, and generate operating condition identification features. The photovoltaic operating conditions are determined based on the operating condition identification characteristics, and operating condition calibration operation data is generated based on the photovoltaic operating conditions and standardized photovoltaic operation data.

4. The photovoltaic array fault diagnosis and location method based on big data and machine learning according to claim 1, characterized in that, The generation of the spatiotemporal multi-scale photovoltaic array topology map includes: Read the acquisition time identifier from the device access relationship and standardized photovoltaic operation data; A multi-scale node set is generated based on the device access relationship, and spatial topology edges are established from string-level nodes to combiner box-level nodes and from combiner box-level nodes to inverter-level nodes under the same acquisition time identifier based on the connection affiliation relationship. Based on the acquisition time identifier, establish a time propagation edge between the same multi-scale nodes corresponding to adjacent acquisition time identifiers; Establish cross-scale propagation edges between string-level nodes and combiner box-level nodes, and between combiner box-level nodes and inverter-level nodes, based on connection affiliation relationships; By associating multi-scale node sets, spatial topological edges, temporal propagation edges, and cross-scale propagation edges according to the acquisition time identifier, a spatiotemporal multi-scale photovoltaic array topology map is generated.

5. The photovoltaic array fault diagnosis and location method based on big data and machine learning according to claim 1, characterized in that, The generation of the multi-scale Gaussian evidence vector and the nodal noise posterior parameters includes: Read the calibration electrical operation field, theoretical benchmark field, and acquisition time identifier from the operating condition calibration operation data, and read the multi-scale node set, spatial topology edge, temporal propagation edge, and cross-scale propagation edge from the spatiotemporal multi-scale photovoltaic array topology graph; The calibration electrical operation fields are matched with the multi-scale node set according to the acquisition time identifier. Based on the calibration electrical operation fields and theoretical benchmark fields corresponding to string-level nodes, combiner box-level nodes and inverter-level nodes, the operation residuals of each topology node are generated. Gaussian evidence vectors for each topology node are generated based on the operational residuals of each topology node. The Gaussian evidence vectors of each topology node are then arranged to generate multi-scale Gaussian evidence vectors. Calculate the fluctuation amplitude of the operating residual of the same topology node under continuous acquisition time markers and the difference in operating residual between adjacent topology nodes, estimate the noise intensity and noise confidence coefficient of each topology node, and generate node noise posterior parameters.

6. The photovoltaic array fault diagnosis and location method based on big data and machine learning according to claim 1, characterized in that, The generation of the node spatiotemporal fault evidence representation includes: The multi-scale Gaussian evidence vector, node noise posterior parameters, and spatiotemporal multi-scale photovoltaic array topology graph are input into the learnable message propagation layer of the photovoltaic array learnable GaBP diagnostic localization model. The learnable message propagation layer includes a propagation edge reading unit, an edge message input construction unit, a learnable Gaussian message generation unit, a noise accuracy correction unit, and a joint posterior update unit. The photovoltaic array learnable GaBP diagnostic localization model includes a learnable message propagation layer, a posterior uncertainty localization output layer, and a fault propagation path backtracking layer. In the learnable message propagation layer, the spatial topology edges, temporal propagation edges, and cross-scale propagation edges in the spatiotemporal multi-scale photovoltaic array topology graph are read by the propagation edge reading unit, and the spatial topology edges, temporal propagation edges, and cross-scale propagation edges are determined as the propagation edge set. The propagation edge set and multi-scale Gaussian evidence vector are input into the side message input construction unit. The corresponding Gaussian evidence components are read according to the start node and end node of each propagation edge, and the side message input vector is generated based on the start node Gaussian evidence component, the end node Gaussian evidence component and the propagation edge type. The learnable Gaussian message generation unit generates Gaussian messages based on the side message input vector; In the noise accuracy correction unit, the corrected message accuracy is generated based on the node noise posterior parameters, and the corrected Gaussian message is generated based on the message mean and the corrected message accuracy. The joint posterior update unit updates the termination node according to the Gaussian message after aggregation and correction of the termination node of the propagation edge, and generates a spatiotemporal fault evidence representation of the node.

7. The photovoltaic array fault diagnosis and location method based on big data and machine learning according to claim 1, characterized in that, The generation of the fault location, location reliability, and initial warning level includes: The spatiotemporal fault evidence representation of the node is input into the posterior uncertainty localization output layer, which includes a candidate localization scoring unit, a fault location and credibility generation unit, and an initial warning level generation unit. In the posterior uncertainty positioning output layer, the spatiotemporal fault evidence representation of the node is read through the candidate positioning scoring unit to generate a candidate positioning node set, and a candidate position score is generated based on each candidate positioning node in the candidate positioning node set. The fault location and credibility generation unit determines the fault location by comparing the candidate location scores of each candidate location node in the candidate location node set, and generates location credibility based on the fault location. The fault location and location reliability are input into the initial warning level generation unit to generate the initial warning level.

8. The photovoltaic array fault diagnosis and location method based on big data and machine learning according to claim 1, characterized in that, The generation of the photovoltaic array fault diagnosis and location results includes: The spatiotemporal fault evidence representation of the node, the fault location, the location credibility, and the initial warning level are input into the fault propagation path backtracking layer. The fault propagation path backtracking layer includes a path starting point determination unit, a path reverse expansion unit, a path propagation intensity calculation unit, and a warning level correction unit. In the fault propagation path backtracking layer, the path starting point determination unit reads the fault location and node spatiotemporal fault evidence representation, and determines the access object number and collection time identifier corresponding to the fault location as the path starting point. The spatiotemporal fault evidence of the node is input into the path backward expansion unit. Starting from the path origin, associated nodes are filtered and a candidate backtracking node sequence is generated. The path propagation intensity calculation unit generates path propagation intensity based on adjacent nodes in the candidate backtracking node sequence, and determines the candidate backtracking node sequence with the highest path propagation intensity as the fault propagation path; The warning level correction unit takes the initial warning level, location reliability, and path propagation strength as inputs, generates a warning level correction amount based on the location reliability and path propagation strength, corrects the initial warning level using the warning level correction amount, generates a corrected warning level, and combines the fault location, location reliability, fault propagation path, and corrected warning level into a photovoltaic array fault diagnosis and location result.