A method and system for statistical analysis of photovoltaic power plant operation data

By dividing the photovoltaic power plant into component clusters, constructing collaborative communication links, and implementing reverse inference mechanisms, the problems of disordered and lost data in photovoltaic power plant statistics have been solved, enabling accurate data statistics and operation and maintenance support.

CN121880422BActive Publication Date: 2026-05-26POWERCHINA JIANGXI ELECTRIC POWER ENGINEERING CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
POWERCHINA JIANGXI ELECTRIC POWER ENGINEERING CO LTD
Filing Date
2026-03-16
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

The existing photovoltaic power plant operation data statistics system has problems such as disordered reordering, overlapping statistical windows, and data loss, which leads to distorted statistical results and affects power plant operation and maintenance decisions and equipment lifespan.

Method used

By dividing the photovoltaic modules into clusters according to their installation areas and series-parallel topology, an ordered subset of data is generated. A data statistics window and a terminal collaborative communication link are constructed. In reverse deduction is performed in conjunction with the physical mechanism of photovoltaic power plant operation to calculate the fit and integrate the final statistical results.

Benefits of technology

It enables accurate and reliable statistics of photovoltaic power plant operation data, ensuring the orderliness and integrity of the data and supporting efficient operation and maintenance decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121880422B_ABST
    Figure CN121880422B_ABST
Patent Text Reader

Abstract

This invention provides a method and system for statistical analysis of photovoltaic power plant operation data. The method includes: dividing the photovoltaic modules into several clusters based on their installation area and series-parallel topology to identify the common data characteristics of each cluster; simultaneously performing hierarchical filtering on the raw data uploaded from each cluster to generate corresponding ordered data subsets; constructing a corresponding data statistics window and simultaneously constructing a corresponding terminal data association matrix; outputting preliminary statistical results based on the data statistics window; simultaneously performing reverse deduction based on the terminal data association matrix according to the operating physical mechanism of the photovoltaic power plant to generate a corresponding deduced data sequence; calculating the degree of fit between the preliminary statistical results and the deduced data sequence; and simultaneously integrating the preliminary statistical results into a corresponding target statistical report when the degree of fit meets preset requirements. This invention can accurately statistically analyze various types of data, thereby improving statistical efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of photovoltaic power plant data processing technology, and in particular to a method and system for statistical analysis of photovoltaic power plant operation data. Background Technology

[0002] Accurate statistical analysis of photovoltaic power plant operation data is crucial for ensuring the smooth operation and maintenance, revenue settlement, and compliance verification of power plants. Due to the dispersed nature of power plant equipment, the complex outdoor environment, and unstable network transmission, existing statistical systems often employ an architecture of "edge acquisition - local caching - fragmented upload - cloud aggregation." Edge-cached data fragmented upload is a critical link connecting the front-end and back-end, directly determining the accuracy and continuity of the statistical results. Currently, existing statistical methods and systems still face significant technical bottlenecks in this area.

[0003] The core flaws of existing technologies lie in the imperfect algorithms related to fragmented uploads: First, the disordered reordering relies solely on a single timestamp without combining multi-dimensional identification information, which can easily lead to data inversion and duplicate misjudgments; second, the time window splicing lacks precise calibration and adaptive processing, which can easily result in overlapping statistical windows or data loss; and third, the breakpoint resume verification mechanism is simple and lacks complete verification and rollback logic, which can easily lead to data loss and duplicate uploads.

[0004] Furthermore, the aforementioned deficiencies directly lead to distorted statistical results, causing inconsistencies between station statistics and grid settlement and meter readings, resulting in revenue disputes and compliance risks. They also mislead operation and maintenance decisions, reduce power plant efficiency, and impact equipment lifespan. Currently, the industry lacks statistical methods and systems that balance reliable transmission, orderly timing, and accurate statistics, necessitating the optimization of existing technologies. Summary of the Invention

[0005] Based on this, the purpose of the present invention is to provide a method and system for statistical analysis of photovoltaic power plant operation data, so as to solve the problems of disordered rearrangement, overlapping statistical windows, and data loss that easily occur in the process of statistical analysis of photovoltaic power plant operation data in the prior art.

[0006] The first aspect of the present invention proposes:

[0007] A method for statistical analysis of photovoltaic power plant operation data, wherein the method includes:

[0008] The photovoltaic modules are divided into several module clusters according to their installation area and series-parallel topology to identify the common data characteristics of each cluster. The raw data uploaded by each cluster is simultaneously processed by hierarchical filtering to generate corresponding ordered data subsets.

[0009] Based on the ordered data subset, the amount of data to be processed and the task priority are detected to construct a corresponding data statistics window. At the same time, a collaborative communication link between power plant monitoring terminals is established to construct a corresponding terminal data association matrix.

[0010] Based on the preliminary statistical results output by the data statistics window, and simultaneously based on the terminal data association matrix, the preliminary statistical results are reverse-engineered according to the operating physical mechanism of the photovoltaic power station to generate the corresponding inference data sequence.

[0011] The degree of fit between the preliminary statistical results and the inferred data sequence is calculated. Simultaneously, when the degree of fit is detected to meet the preset requirements, the preliminary statistical results are integrated into the corresponding target statistical report.

[0012] The beneficial effects of this invention are as follows: This technical solution can effectively solve the problems of disordered rearrangement, overlapping statistical windows, and data loss in the statistical analysis of existing photovoltaic power plant operation data: By dividing clusters according to component installation areas and topological relationships and filtering data in a hierarchical manner, the orderliness of data is ensured; by constructing statistical windows as needed, building terminal collaborative communication links and correlation matrices, the problem of window overlap is avoided; by combining the power plant operation mechanism with reverse deduction to verify data consistency, the risk of data loss is reduced, and finally, accurate and reliable statistics of power plant operation data are achieved, providing strong data support for efficient operation and maintenance.

[0013] Furthermore, the step of detecting the amount of data to be processed and the task priority based on the ordered data subset to construct a corresponding data statistics window includes:

[0014] The short-term fluctuation correlation and long-term trend correlation in the ordered data subset are mined by the temporal attention mechanism to generate corresponding temporal effectiveness weights, and the amount of data to be processed is extracted from the ordered data subset according to the temporal effectiveness weights.

[0015] The amount of data to be processed is associated with the fault prediction risk level to generate the corresponding task priority. Simultaneously, based on the serial-parallel topology, the amount of data to be processed and the task priority are merged into a window construction unit.

[0016] Constrained by the bandwidth carrying capacity of the photovoltaic power station, the data sampling frequency and data buffer capacity of the window construction unit are dynamically matched, and the topology branch loss compensation coefficient is embedded synchronously to construct the data statistics window accordingly.

[0017] Furthermore, the step of establishing a collaborative communication link between power plant monitoring terminals to construct a corresponding terminal data association matrix includes:

[0018] Collect the actual temperature and voltage difference data of the photovoltaic modules corresponding to each of the power station monitoring terminals to calculate the corresponding transmission risk level, and simultaneously divide the power station monitoring terminals with similar transmission risk levels and adjacent physical locations into collaborative communication units.

[0019] Receive target monitoring data sent by each of the aforementioned collaborative communication units, and synchronously verify the integrity of the target monitoring data through hash value verification.

[0020] If the target monitoring data is found to be complete, the corresponding core association features are extracted, and each core association feature is simultaneously processed into a matrix using a sparse matrix reconstruction algorithm to generate the corresponding terminal data association matrix.

[0021] Furthermore, the step of reverse-engineering the preliminary statistical results based on the terminal data association matrix according to the operating physical mechanism of the photovoltaic power station to generate the corresponding deduced data sequence includes:

[0022] The operating physical mechanism of the photovoltaic power station is transformed into hard boundary constraints for reverse deduction. At the same time, topological features strongly correlated with the hard boundary constraints are extracted from the terminal data association matrix to construct a two-dimensional deduction baseline with mechanism constraints and topological features.

[0023] Using the abnormal data nodes in the preliminary statistical results as the starting point for tracing the source, and combining the dual-dimensional inference baseline, the historical operation data and interaction relationships corresponding to the starting point are traced through the terminal data association matrix. Combined with the operation physical mechanism of the photovoltaic power station, the source cause and propagation path of the data anomaly are deduced in reverse, and the corresponding tracing and inference path is generated.

[0024] The source tracing and deduction path is dynamically calibrated to output the corresponding deduction data sequence.

[0025] Furthermore, the step of dynamically calibrating the source tracing and deduction path to output the deduction data sequence includes:

[0026] The operating physical mechanism of the photovoltaic power station is transformed into a corresponding inverse mapping function, and the topological correlation of the terminal data correlation matrix is ​​combined simultaneously to select effective nodes in the source tracing and deduction path.

[0027] Extract the temporal operation characteristics of the effective nodes, and simultaneously match the temporal anchoring nodes corresponding to the effective nodes in the terminal data association matrix according to the temporal operation characteristics;

[0028] The timing and numerical deviations of the effective nodes are dynamically corrected using the timing reference data of the timing anchor nodes to generate corresponding target nodes. Simultaneously, the target nodes are integrated and processed to generate the inference data sequence.

[0029] Furthermore, the step of calculating the degree of fit between the preliminary statistical results and the inferred data sequence includes:

[0030] The preliminary statistical results are traced back to the ordered subset of data, and the inferred data sequence is simultaneously traced back to the terminal data association matrix to construct a dual-end tracing link.

[0031] By comparing the core data features in the dual-end traceability link, cross-verification is performed simultaneously to eliminate abnormal deviation data without traceability support, and the corresponding initial fit is calculated.

[0032] The initial fit is dynamically corrected to generate the corresponding fit.

[0033] Furthermore, the step of dynamically correcting the initial fit to generate the corresponding fit includes:

[0034] Based on the terminal data association matrix, the implicit association factors between the various power station monitoring terminals are mined, and the corresponding coupling correction vector is created by combining the core parameters of the operating physical mechanism of the photovoltaic power station.

[0035] A correction benchmark is generated based on the characteristics of the same source data to match the initial fit, and the corresponding correction coefficient is matched simultaneously based on the magnitude of the correction benchmark.

[0036] Based on the correction coefficient and the coupling correction vector, a corresponding target correction vector is generated, and the target correction vector is simultaneously fused with the initial fit to generate the corresponding fit.

[0037] The second aspect of the present invention proposes:

[0038] A photovoltaic power plant operation data statistics system, wherein the system includes:

[0039] The identification module is used to divide the photovoltaic modules into several module clusters according to their installation area and series-parallel topology, so as to identify the common data characteristics of each cluster and simultaneously perform hierarchical filtering processing on the raw data uploaded by each cluster to generate corresponding ordered data subsets.

[0040] The construction module is used to detect the amount of data to be processed and the task priority based on the ordered data subset, so as to construct the corresponding data statistics window and simultaneously build the collaborative communication link between the power plant monitoring terminals to construct the corresponding terminal data association matrix.

[0041] The deduction module is used to output the corresponding preliminary statistical results according to the data statistics window, and simultaneously, based on the terminal data association matrix, to reverse-deduce the preliminary statistical results according to the operating physical mechanism of the photovoltaic power station, so as to generate the corresponding deduction data sequence.

[0042] An integration module is used to calculate the degree of fit between the preliminary statistical results and the inferred data sequence, and simultaneously integrate the preliminary statistical results into a corresponding target statistical report when the degree of fit is detected to meet preset requirements.

[0043] Furthermore, the building module is specifically used for:

[0044] The short-term fluctuation correlation and long-term trend correlation in the ordered data subset are mined by the temporal attention mechanism to generate corresponding temporal effectiveness weights, and the amount of data to be processed is extracted from the ordered data subset according to the temporal effectiveness weights.

[0045] The amount of data to be processed is associated with the fault prediction risk level to generate the corresponding task priority. Simultaneously, based on the serial-parallel topology, the amount of data to be processed and the task priority are merged into a window construction unit.

[0046] Constrained by the bandwidth carrying capacity of the photovoltaic power station, the data sampling frequency and data buffer capacity of the window construction unit are dynamically matched, and the topology branch loss compensation coefficient is embedded synchronously to construct the data statistics window accordingly.

[0047] Furthermore, the building module is specifically used for:

[0048] Collect the actual temperature and voltage difference data of the photovoltaic modules corresponding to each of the power station monitoring terminals to calculate the corresponding transmission risk level, and simultaneously divide the power station monitoring terminals with similar transmission risk levels and adjacent physical locations into collaborative communication units.

[0049] Receive target monitoring data sent by each of the aforementioned collaborative communication units, and synchronously verify the integrity of the target monitoring data through hash value verification.

[0050] If the target monitoring data is found to be complete, the corresponding core association features are extracted, and each core association feature is simultaneously processed into a matrix using a sparse matrix reconstruction algorithm to generate the corresponding terminal data association matrix.

[0051] Furthermore, the inference module is specifically used for:

[0052] The operating physical mechanism of the photovoltaic power station is transformed into hard boundary constraints for reverse deduction. At the same time, topological features strongly correlated with the hard boundary constraints are extracted from the terminal data association matrix to construct a two-dimensional deduction baseline with mechanism constraints and topological features.

[0053] Using the abnormal data nodes in the preliminary statistical results as the starting point for tracing the source, and combining the dual-dimensional inference baseline, the historical operation data and interaction relationships corresponding to the starting point are traced through the terminal data association matrix. Combined with the operation physical mechanism of the photovoltaic power station, the source cause and propagation path of the data anomaly are deduced in reverse, and the corresponding tracing and inference path is generated.

[0054] The source tracing and deduction path is dynamically calibrated to output the corresponding deduction data sequence.

[0055] Furthermore, the inference module is specifically used for:

[0056] The operating physical mechanism of the photovoltaic power station is transformed into a corresponding inverse mapping function, and the topological correlation of the terminal data correlation matrix is ​​combined simultaneously to select effective nodes in the source tracing and deduction path.

[0057] Extract the temporal operation characteristics of the effective nodes, and simultaneously match the temporal anchoring nodes corresponding to the effective nodes in the terminal data association matrix according to the temporal operation characteristics;

[0058] The timing and numerical deviations of the effective nodes are dynamically corrected using the timing reference data of the timing anchor nodes to generate corresponding target nodes. Simultaneously, the target nodes are integrated and processed to generate the inference data sequence.

[0059] Furthermore, the integration module is specifically used for:

[0060] The preliminary statistical results are traced back to the ordered subset of data, and the inferred data sequence is simultaneously traced back to the terminal data association matrix to construct a dual-end tracing link.

[0061] By comparing the core data features in the dual-end traceability link, cross-verification is performed simultaneously to eliminate abnormal deviation data without traceability support, and the corresponding initial fit is calculated.

[0062] The initial fit is dynamically corrected to generate the corresponding fit.

[0063] Furthermore, the integration module is specifically used for:

[0064] Based on the terminal data association matrix, the implicit association factors between the various power station monitoring terminals are mined, and the corresponding coupling correction vector is created by combining the core parameters of the operating physical mechanism of the photovoltaic power station.

[0065] A correction benchmark is generated based on the characteristics of the same source data to match the initial fit, and the corresponding correction coefficient is matched simultaneously based on the magnitude of the correction benchmark.

[0066] Based on the correction coefficient and the coupling correction vector, a corresponding target correction vector is generated, and the target correction vector is simultaneously fused with the initial fit to generate the corresponding fit.

[0067] The third aspect of the present invention proposes:

[0068] A computer includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the photovoltaic power plant operation data statistics method as described above.

[0069] The fourth aspect of the present invention proposes:

[0070] A readable storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the photovoltaic power plant operation data statistics method as described above.

[0071] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0072] Figure 1 A flowchart of a photovoltaic power plant operation data statistics method provided in the first embodiment of the present invention;

[0073] Figure 2 The structural block diagram of the photovoltaic power plant operation data statistics system provided in the third embodiment of the present invention.

[0074] The following detailed description, in conjunction with the accompanying drawings, will further illustrate the present invention. Detailed Implementation

[0075] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Several embodiments of the invention are illustrated in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete.

[0076] It should be noted that when a component is said to be "fixed to" another component, it can be directly on the other component or there may be an intervening component. When a component is said to be "connected to" another component, it can be directly connected to the other component or there may be an intervening component. The terms "vertical," "horizontal," "left," "right," and similar expressions used in this document are for illustrative purposes only.

[0077] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.

[0078] Please see Figure 1 The image shows a photovoltaic power plant operation data statistics method provided in the first embodiment of the present invention. The photovoltaic power plant operation data statistics method provided in this embodiment can accurately and completely collect various types of data, thereby improving the data statistics efficiency.

[0079] Specifically, this embodiment provides:

[0080] A method for statistical analysis of photovoltaic power plant operation data, wherein the method includes:

[0081] Step S10: Divide the photovoltaic modules into several module clusters according to the installation area and series-parallel topology to identify the common data characteristics of each cluster, and simultaneously perform hierarchical filtering processing on the raw data uploaded by each cluster to generate corresponding ordered data subsets.

[0082] It is important to note that, firstly, photovoltaic power plants typically contain a large number of photovoltaic modules, which are distributed by region and connected in series and parallel topologies. Module data from different regions and different topological branches exhibit commonalities (e.g., current data from modules in the same series are correlated). Traditional data statistics do not perform targeted classification, easily leading to data confusion and redundancy. Therefore, the photovoltaic modules are divided into several clusters based on their installation area (e.g., east, west, and north arrays) and series-parallel topology (e.g., a specific series or grid-connected branch). Cluster division identifies common data characteristics within each cluster (e.g., consistent voltage fluctuation patterns and power output trends among modules in the same cluster). Simultaneously, the raw data uploaded from each cluster (e.g., module current, voltage, power, temperature, combiner box and inverter operation data) undergoes hierarchical filtering (first-level filtering removes obviously invalid data such as sensor fault data; second-level filtering removes redundant data such as duplicated data; third-level filtering extracts core data such as key operating parameters), generating corresponding ordered data subsets. Specifically, cluster division ensures clear data classification, and hierarchical filtering improves data quality, laying the foundation for subsequent statistical processing.

[0083] Step S20: Detect the amount of data to be processed and the task priority based on the ordered data subset to construct a corresponding data statistics window, and simultaneously build a collaborative communication link between power plant monitoring terminals to construct a corresponding terminal data association matrix.

[0084] It should be noted that, secondly, the amount of data to be processed in a photovoltaic power station varies depending on the time period and operating state (e.g., the amount of data during periods of sufficient sunlight is greater than that during cloudy days) and the task priority (e.g., the priority of fault warning-related data statistics is higher than that of regular operating data). Traditional fixed statistical models cannot adapt to this dynamic characteristic. Therefore, based on ordered data subsets, the amount of data to be processed (statistically calculating the current effective data scale to be processed) and the task priority (priority is determined by factors such as fault prediction, operation and maintenance requirements, and grid connection assessment, with fault-related data having the highest priority and regular operating data having the next highest priority) are constructed. Based on the amount of data to be processed and the task priority, a corresponding data statistics window is constructed (the window includes core parameters such as data sampling frequency, cache capacity, and statistical period, such as high-priority tasks...). The system should have short statistical periods and high sampling frequencies. Meanwhile, photovoltaic power plants contain multiple monitoring terminals (such as module-level sensors, combiner box monitoring terminals, inverter monitoring terminals, and grid-connected cabinet monitoring terminals). Since the data from each terminal is dispersed, a collaborative communication link needs to be established to achieve data interaction. Therefore, a collaborative communication link (such as a dedicated communication link based on 5G or LoRa to ensure real-time data transmission between terminals) should be built between the power plant monitoring terminals. Corresponding data from each terminal is collected through this link, and a corresponding terminal data correlation matrix is ​​constructed (matrix elements represent the correlation strength of different terminal data, such as the correlation strength between the current data of a component sensor and the corresponding combiner box terminal). Specifically, the data statistics window dynamically adapts, and the collaborative communication link and correlation matrix break down the barriers between terminal data, ensuring the real-time performance and correlation of statistical processing.

[0085] Step S30: Output the corresponding preliminary statistical results according to the data statistics window, and simultaneously, based on the terminal data association matrix, reverse the preliminary statistical results according to the operating physical mechanism of the photovoltaic power station to generate the corresponding deduced data sequence.

[0086] It should be noted that, next, preliminary statistics are performed on the ordered data subset through the data statistics window (such as the average power, maximum voltage, and frequency of faults of a certain cluster component), and the corresponding preliminary statistical results are output. However, the preliminary statistical results are only based on the surface characteristics of the data and do not take into account the physical mechanism of photovoltaic power plant operation (such as the correlation mechanism between photovoltaic module power and light and temperature, and the power loss mechanism of series and parallel topologies), which may result in statistical bias. Therefore, based on the terminal data correlation matrix (which clarifies the correlation relationship of terminal data), the preliminary statistical results are reverse-engineered according to the physical mechanism of photovoltaic power plant operation (such as, based on the preliminary statistics of abnormal power in a certain branch, reverse-engineering whether the light, temperature, and topology connection status of the components in that branch conform to physical laws), generating the corresponding deduced data sequence (such as the power and voltage data sequence that the branch should have based on the physical mechanism). Specifically, the reverse deduction verifies the rationality of the preliminary statistical results through the physical mechanism, making up for the limitations of pure data statistics.

[0087] Step S40: Calculate the degree of fit between the preliminary statistical results and the inferred data sequence, and simultaneously integrate the preliminary statistical results into the corresponding target statistical report when the degree of fit is detected to meet the preset requirements.

[0088] It should be noted that, finally, to ensure the accuracy of the statistical results, the degree of fit between the preliminary statistical results and the inferred data sequence is calculated (quantifying the consistency between the two; a higher degree of fit indicates that the preliminary statistical results are more in line with physical laws). Simultaneously, a preset requirement for the degree of fit is set (e.g., a degree of fit ≥90% is considered satisfactory, which can be adjusted according to power plant operation and maintenance standards). When the degree of fit meets the preset requirement, it indicates that the preliminary statistical results are true and reliable, and they are integrated into the corresponding target statistical report (the report includes core content such as statistics on the operating parameters of each cluster component, terminal operating status, fault statistics, and grid connection performance statistics), ultimately completing the accurate statistical analysis of photovoltaic power plant operation data. If the degree of fit does not meet the requirements, the previous steps are returned for reprocessing (e.g., re-filtering data, adjusting the statistical window, and correcting the reverse inference parameters), forming a closed-loop optimization.

[0089] Second Embodiment

[0090] Furthermore, the step of detecting the amount of data to be processed and the task priority based on the ordered data subset to construct a corresponding data statistics window includes:

[0091] The short-term fluctuation correlation and long-term trend correlation in the ordered data subset are mined by the temporal attention mechanism to generate corresponding temporal effectiveness weights, and the amount of data to be processed is extracted from the ordered data subset according to the temporal effectiveness weights.

[0092] The amount of data to be processed is associated with the fault prediction risk level to generate the corresponding task priority. Simultaneously, based on the serial-parallel topology, the amount of data to be processed and the task priority are merged into a window construction unit.

[0093] Constrained by the bandwidth carrying capacity of the photovoltaic power station, the data sampling frequency and data buffer capacity of the window construction unit are dynamically matched, and the topology branch loss compensation coefficient is embedded synchronously to construct the data statistics window accordingly.

[0094] It is important to note that, firstly, photovoltaic power plant operation data exhibits significant time-series characteristics, including short-term fluctuations (such as short-term power fluctuations caused by cloud cover) and long-term trends (such as overall power trends caused by intraday changes in sunlight and annual power trends caused by seasonal variations). Different time-series characteristics place different demands on the statistical window. Therefore, a time-series attention mechanism (focusing on core data features at different time-series nodes and assigning different attention weights to short-term fluctuations and long-term trends) is used to mine short-term fluctuation correlations (such as the correlation of component power fluctuations within the same hour) and long-term trend correlations (such as the correlation of component power changes within a day) within ordered data subsets, generating corresponding time-series validity weights (such as validity weights for short-term fluctuation data and validity weights for long-term trend data). Simultaneously, based on these time-series validity weights, the amount of data to be processed is extracted from the ordered data subsets (prioritizing the extraction of core data with high validity weights to ensure the core nature and relevance of the data to be processed). Specifically, time-series feature mining ensures accurate extraction of the amount of data to be processed, providing reasonable input for window construction.

[0095] Task prioritization needs to be combined with the core needs of photovoltaic power plants. Fault prediction is the key to power plant operation and maintenance. The amount of data to be processed is directly related to the fault prediction risk (e.g., if the proportion of abnormal data in the data to be processed of a certain cluster is high, the corresponding fault prediction risk is high). Therefore, the amount of data to be processed is associated with the fault prediction risk (the risk is calculated by the fault prediction model; the higher the proportion of abnormal data in the amount of data to be processed, the higher the risk), and corresponding task priorities are generated (e.g., high fault prediction risk corresponds to high task priority, and low risk corresponds to low task priority). Simultaneously, based on the series and parallel topology relationship (the component data of different topology branches are strongly correlated, and the statistical tasks of the same topology branch need to be processed collaboratively), the amount of data to be processed and the task priority are grouped into window building units (e.g., the data to be processed and the corresponding task priority of the same series of branches are grouped into one window building unit to ensure the continuity of data statistics of the same topology branch). Specifically, the task priority is grouped and the window building unit is made to fit the topology characteristics to improve the window adaptability.

[0096] The construction of the data statistics window needs to take into account both communication bandwidth constraints and topology characteristics. The communication bandwidth carrying capacity of photovoltaic power plants is limited (such as the upper limit of terminal data transmission bandwidth). If the window sampling frequency is too high and the buffer capacity is too large, it will lead to bandwidth congestion. Therefore, the bandwidth carrying capacity of the photovoltaic power plant is used as a constraint (the maximum sampling frequency and buffer capacity of the window are determined according to the upper limit of bandwidth). The data sampling frequency (high priority tasks and large data volume correspond to high sampling frequency, and low priority tasks and small data volume correspond to low sampling frequency) and data buffer capacity (the larger the data volume and the longer the statistical period, the larger the buffer capacity) of the window construction unit are dynamically matched. At the same time, there are branch losses (such as line losses and contact losses) in the series and parallel topology of photovoltaic modules. Losses will affect the accuracy of data statistics. Therefore, the topology branch loss compensation coefficient is embedded (the loss coefficient is calculated according to the topology branch length, material and other parameters). The loss compensation logic is incorporated into the window construction (such as reserving loss compensation space in advance when collecting data). Finally, the corresponding data statistics window is constructed. Specifically, the bandwidth constraint and loss compensation ensure that the statistics window is both adapted to the communication capacity and conforms to the topology loss characteristics, ensuring the accuracy and real-time performance of the subsequent statistical results.

[0097] Furthermore, the step of establishing a collaborative communication link between power plant monitoring terminals to construct a corresponding terminal data association matrix includes:

[0098] Collect the actual temperature and voltage difference data of the photovoltaic modules corresponding to each of the power station monitoring terminals to calculate the corresponding transmission risk level, and simultaneously divide the power station monitoring terminals with similar transmission risk levels and adjacent physical locations into collaborative communication units.

[0099] Receive target monitoring data sent by each of the aforementioned collaborative communication units, and synchronously verify the integrity of the target monitoring data through hash value verification.

[0100] If the target monitoring data is found to be complete, the corresponding core association features are extracted, and each core association feature is simultaneously processed into a matrix using a sparse matrix reconstruction algorithm to generate the corresponding terminal data association matrix.

[0101] It should be noted that, firstly, photovoltaic power plant monitoring terminals are widely distributed, and the operating environments (e.g., high surface temperature of the terminal on the module, high humidity inside the combiner box) and transmission conditions (e.g., weak signal of the terminal in the edge area) of different terminals vary, resulting in different transmission risks (e.g., high data transmission failure rate of the terminal in high temperature environment). Therefore, the actual temperature of the photovoltaic module (the surface temperature of the module where the terminal is located) and voltage difference data (the difference between the input and output voltage of the terminal, which may often indicate terminal failure) of each power plant monitoring terminal are collected. By judging the temperature threshold (e.g., temperature > 60℃, transmission risk increases) and analyzing the voltage difference (e.g., voltage difference > 5V, transmission risk increases), the corresponding transmission risk level (e.g., high, medium, and low levels) is calculated. At the same time, power plant monitoring terminals with similar transmission risk levels and adjacent physical locations are divided into collaborative communication units (module sensor terminals in the same area and with the same risk level are grouped into one unit). Specifically, the division into collaborative communication units can improve communication stability, avoid the impact of a single terminal failure on overall communication, and facilitate centralized management.

[0102] Each collaborative communication unit sends target monitoring data (such as operating parameters and fault information of each terminal within the unit) through the collaborative communication link. During transmission, the data may be lost or tampered with due to signal interference or link interruption, affecting the accuracy of subsequent statistics. Therefore, after receiving the target monitoring data sent by each collaborative communication unit, the integrity of the data is confirmed by hash value verification (the hash value of the original data is calculated, the hash value of the received data is calculated at the receiving end, and the two are compared to see if they are consistent. If they are consistent, the data is complete). Specifically, hash value verification ensures that the transmitted data has not been tampered with or lost, thus guaranteeing data quality.

[0103] If the target monitoring data is found to be complete, it indicates that the data can be used for subsequent processing. At this time, the corresponding core correlation features are extracted (such as the current correlation features of terminals within the same collaborative unit, the fault correlation features of different terminals, and the correlation features between component data and inverter data). Since there are many monitoring terminals in a photovoltaic power station and the core correlation features have high dimensionality, they need to be matrixed to clearly present the correlation relationship. Therefore, a sparse matrix reconstruction algorithm (suitable for processing high-dimensional and sparse data, which can remove invalid correlations and retain core correlations) is used to matrixize each core correlation feature (the matrix rows and columns correspond to different monitoring terminals, and the element values ​​represent the correlation strength between terminals). Finally, a terminal data correlation matrix is ​​generated. Specifically, this matrix intuitively presents the correlation relationship of the data of each monitoring terminal, providing a core correlation basis for reverse inference and fit calculation.

[0104] Furthermore, the step of reverse-engineering the preliminary statistical results based on the terminal data association matrix according to the operating physical mechanism of the photovoltaic power station to generate the corresponding deduced data sequence includes:

[0105] The operating physical mechanism of the photovoltaic power station is transformed into hard boundary constraints for reverse deduction. At the same time, topological features strongly correlated with the hard boundary constraints are extracted from the terminal data association matrix to construct a two-dimensional deduction baseline with mechanism constraints and topological features.

[0106] Using the abnormal data nodes in the preliminary statistical results as the starting point for tracing the source, and combining the dual-dimensional inference baseline, the historical operation data and interaction relationships corresponding to the starting point are traced through the terminal data association matrix. Combined with the operation physical mechanism of the photovoltaic power station, the source cause and propagation path of the data anomaly are deduced in reverse, and the corresponding tracing and inference path is generated.

[0107] The source tracing and deduction path is dynamically calibrated to output the corresponding deduction data sequence.

[0108] It should be noted that, firstly, the operation of photovoltaic power plants follows clear physical mechanisms (such as the photovoltaic module power formula P=UI, Kirchhoff's laws for series and parallel circuits, and the correlation mechanism between module power and irradiance G and temperature T: P=Pn×[1+α(T-Tn)]×(G / Gn), where Pn is the power under standard test conditions, α is the power temperature coefficient, Tn is the standard test temperature, and Gn is the standard test irradiance). Reverse deduction must be constrained by these physical mechanisms to avoid deviations from reality. Therefore, the operating physical mechanisms of photovoltaic power plants are transformed into… The hard boundary constraints of the reverse deduction (such as voltage and current needing to meet the laws of series and parallel circuits, and power needing to conform to the correlation laws of illumination and temperature) are used to simultaneously extract topological features strongly correlated with the hard boundary constraints from the terminal data correlation matrix (such as equal current for components in the same series circuit, equal voltage at the same grid connection point, and correlation features between topological branch losses and current). This constructs a two-dimensional deduction baseline with mechanistic constraints and topological features. Specifically, the two-dimensional deduction baseline ensures that the reverse deduction conforms to physical laws and fits the actual topology of the power station, providing a reliable benchmark for the deduction.

[0109] Preliminary statistical results may contain anomalous data nodes (e.g., a component's power is significantly lower than other components in the same cluster, or a branch current fluctuates abnormally). These anomalous nodes are the core targets for reverse engineering. Therefore, the anomalous data nodes in the preliminary statistical results are taken as the starting point for tracing (e.g., a component with abnormal power). Combining the dual-dimensional inference baseline (clarifying physical laws and topological relationships), the historical operating data (e.g., the component's power and temperature data over the past 24 hours) and interaction relationships (e.g., the data flow interaction relationship between the component and the combiner box and inverter, and the interaction relationship with other components in the same circuit) are traced through the terminal data association matrix to the starting point. The correlation between components is further analyzed by combining the physical mechanisms of photovoltaic power plant operation (such as power anomalies possibly originating from shading, component aging, excessive temperature, or topology connection failures). The source cause of data anomalies (such as power anomalies originating from component surface shading) and propagation path (such as shading causing a decrease in component power, which in turn affects the current output of the entire series circuit) are deduced in reverse. Corresponding source deduction paths are generated (such as "component shading → component power anomaly → series current decrease → combiner box data anomaly"). Specifically, the source deduction path clearly presents the cause and propagation logic of abnormal data, providing a core basis for generating deduction data sequences.

[0110] The source tracing and extrapolation path may contain deviations (such as those caused by neglecting environmental interference or sensor accuracy errors). Dynamic calibration is required to improve the accuracy of the extrapolation. Therefore, the source tracing and extrapolation path is dynamically calibrated (e.g., calibrating extrapolation parameters by combining historical normal operation data and correcting extrapolation deviations based on sensor accuracy). After calibration, the corresponding extrapolation data sequence is output (e.g., based on the physical mechanism and source tracing path, deriving the power and voltage data sequence that the component should have under normal conditions, as well as the evolution data sequence under abnormal conditions). Specifically, the extrapolation data sequence provides a benchmark for the fit calculation, ensuring the rationality of the statistical results.

[0111] Furthermore, the step of dynamically calibrating the source tracing and deduction path to output the deduction data sequence includes:

[0112] The operating physical mechanism of the photovoltaic power station is transformed into a corresponding inverse mapping function, and the topological correlation of the terminal data correlation matrix is ​​combined simultaneously to select effective nodes in the source tracing and deduction path.

[0113] Extract the temporal operation characteristics of the effective nodes, and simultaneously match the temporal anchoring nodes corresponding to the effective nodes in the terminal data association matrix according to the temporal operation characteristics;

[0114] The timing and numerical deviations of the effective nodes are dynamically corrected using the timing reference data of the timing anchor nodes to generate corresponding target nodes. Simultaneously, the target nodes are integrated and processed to generate the inference data sequence.

[0115] It should be noted that, firstly, to ensure that the calibration is supported by a clear physical mechanism, the operating physical mechanism of the photovoltaic power station is transformed into a corresponding inverse mapping function (e.g., mapping power anomalies to functions caused by factors such as illumination, temperature, and topology faults, and mapping voltage anomalies to functions caused by factors such as series-parallel connections and module aging). Simultaneously, the topological correlation of the terminal data correlation matrix is ​​combined (clarifying the topological connection relationship of each node, such as the relationship between a node and its upstream and downstream nodes), and effective nodes are selected in the source tracing and deduction path (removing redundant nodes unrelated to the cause of the anomaly, such as redundant nodes in other branch terminal nodes in the same area when a module has a power anomaly). Specifically, the inverse mapping function provides the mechanistic basis for calibration, and the selection of effective nodes improves calibration efficiency and accuracy.

[0116] Photovoltaic power plant operation data has strong time-series characteristics, and the operating parameters of different time-series nodes differ. The time-series characteristics of effective nodes in the tracing and extrapolation path (such as data collection timestamps and time-series fluctuation patterns) need to be matched with historical normal data to ensure accurate calibration. Therefore, the time-series operation characteristics of effective nodes (such as hourly power fluctuation patterns and daily power trend characteristics of a certain effective node) are extracted. Simultaneously, based on these time-series operation characteristics, the time-series anchor nodes corresponding to the effective nodes are matched in the terminal data association matrix (such as normal operating nodes under the same historical period and operating conditions, or normal nodes with consistent time-series characteristics within the same cluster). Specifically, the time-series anchor nodes provide a time-series benchmark for the deviation correction of effective nodes.

[0117] The time-series reference data of the anchor nodes (such as power and voltage data during normal operation, and time-series fluctuation range) is the core basis for correcting the deviation of effective nodes. Therefore, by using the time-series reference data of the anchor nodes, the time-series deviation of effective nodes (such as adjusting the data acquisition timestamp to ensure alignment with the reference time series) and the numerical deviation (such as correcting the power and voltage values ​​of effective nodes to conform to the normal time-series fluctuation range) are dynamically corrected. After correction, the corresponding target nodes (nodes whose time series and values ​​conform to physical laws and normal operation characteristics) are generated. Simultaneously, the target nodes are integrated according to time sequence and topological relationship (such as integration by operating period and integration by series and parallel branches) to finally generate the extrapolated data sequence. Specifically, this sequence is precisely calibrated to ensure a high degree of consistency with the actual operation of the photovoltaic power station, providing a reliable comparison basis for subsequent consistency calculations.

[0118] Furthermore, the step of calculating the degree of fit between the preliminary statistical results and the inferred data sequence includes:

[0119] The preliminary statistical results are traced back to the ordered subset of data, and the inferred data sequence is simultaneously traced back to the terminal data association matrix to construct a dual-end tracing link.

[0120] By comparing the core data features in the dual-end traceability link, cross-verification is performed simultaneously to eliminate abnormal deviation data without traceability support, and the corresponding initial fit is calculated.

[0121] The initial fit is dynamically corrected to generate the corresponding fit.

[0122] It should be noted that, firstly, to ensure the traceability of the fit calculation and avoid unfounded biased comparisons, the preliminary statistical results are traced back to the internal of the ordered data subset (clarifying the original data source of the preliminary statistical results, such as which cluster and which time period the ordered data of a certain statistical indicator comes from), and simultaneously the inferred data sequence is traced back to the internal of the terminal data association matrix (clarifying the terminal association source of the inferred data, such as which monitoring terminals the association data of a certain inferred value comes from). Correspondingly, a dual-end tracing link is constructed (the preliminary statistical end tracing link and the inferred data end tracing link). Specifically, the dual-end tracing link ensures that the comparison between the two has clear data source support and avoids invalid comparisons.

[0123] By comparing core data features in the dual-end traceability chain (such as cluster consistency of data sources, consistency of time-series features, and consistency of numerical fluctuation patterns), cross-verification is performed simultaneously (e.g., using the original data of the preliminary statistical results to verify the rationality of the inferred data, and using the inferred data to verify the accuracy of the preliminary statistical results). During the comparison and cross-verification process, abnormal deviation data without traceability support are removed (e.g., abnormal values ​​in the preliminary statistical results without original data support, and deviation values ​​in the inferred data without terminal association support). These data usually originate from sensor failures, data transmission errors, etc., and have no reference value. Based on the data after removing abnormal deviations, the corresponding initial fit is calculated (e.g., using algorithms such as cosine similarity and mean square error to quantify consistency; the closer the cosine similarity is to 1 and the closer the mean square error is to 0, the higher the initial fit). Specifically, core feature comparison and anomaly removal ensure that the initial fit can reflect the true level of consistency.

[0124] The initial fit did not take into account potential hidden influencing factors (such as environmental interference and topology loss fluctuations), and dynamic correction is required to improve accuracy. Therefore, the initial fit is dynamically corrected (e.g., by combining environmental factors and topology loss to correct the fit value), and finally a fit is generated. Specifically, the corrected fit can more realistically and accurately reflect the consistency between the preliminary statistical results and the inferred data sequence, providing a reliable basis for the generation of the target statistical report.

[0125] Furthermore, the step of dynamically correcting the initial fit to generate the corresponding fit includes:

[0126] Based on the terminal data association matrix, the implicit association factors between the various power station monitoring terminals are mined, and the corresponding coupling correction vector is created by combining the core parameters of the operating physical mechanism of the photovoltaic power station.

[0127] A correction benchmark is generated based on the characteristics of the same source data to match the initial fit, and the corresponding correction coefficient is matched simultaneously based on the magnitude of the correction benchmark.

[0128] Based on the correction coefficient and the coupling correction vector, a corresponding target correction vector is generated, and the target correction vector is simultaneously fused with the initial fit to generate the corresponding fit.

[0129] It should be noted that, firstly, in addition to explicit correlations (such as series-parallel topology correlations and current and voltage data correlations) between photovoltaic power plant monitoring terminals, there are also implicit correlation factors (such as temperature correlations between terminals in different areas, power correlations between components under the same inverter, and correlations between environmental sensor and component operation data). These implicit factors affect the accuracy of the fit. Therefore, based on the terminal data correlation matrix, implicit correlation factors between various power plant monitoring terminals are mined (e.g., implicit correlations are extracted through correlation rule mining algorithms). Combining the core parameters of the photovoltaic power plant's operating physical mechanism (such as light intensity, component temperature coefficient, topology loss coefficient, and inverter conversion efficiency), the implicit correlation factors are integrated with the core parameters to create a corresponding coupling correction vector (vector elements represent the correction weights of each implicit correlation factor and core parameter on the fit). Specifically, the coupling correction vector covers both explicit and implicit influencing factors, providing comprehensive support for correction.

[0130] The rationality of the correction benchmark directly determines the correction accuracy. Data characteristics from the same source (such as data characteristics of components in the same cluster or branch of the same topology) are consistent and can serve as the core basis for the correction benchmark. Therefore, a correction benchmark adapted to the initial fit is generated based on the characteristics of the data from the same source (such as the fit range of the same cluster under normal operating conditions or the standard fit benchmark of the same type of photovoltaic power station). At the same time, the corresponding correction coefficient is matched according to the size of the correction benchmark (such as the larger the initial fit deviates from the correction benchmark, the larger the correction coefficient, ensuring that the larger the deviation, the more significant the correction). Specifically, the matching of the correction benchmark and the correction coefficient ensures that the correction is targeted and avoids blind correction.

[0131] Based on the correction coefficients and the coupled correction vectors, a corresponding target correction vector is generated through vector weighting (the coupled correction vectors are weighted according to the correction coefficients to highlight the correction effect of the core influencing factors). Simultaneously, the target correction vector is fused with the initial fit (e.g., the target correction vector is integrated into the initial fit using a vector dot product algorithm, and the initial fit value is adjusted) to generate the final fit. Specifically, the fusion process ensures accurate correction, so that the final fit can truly reflect the consistency between the preliminary statistical results and the inferred data sequence, providing an accurate basis for determining whether to output the target statistical report.

[0132] Please see Figure 2 The third embodiment of the present invention provides:

[0133] A photovoltaic power plant operation data statistics system, wherein the system includes:

[0134] The identification module is used to divide the photovoltaic modules into several module clusters according to their installation area and series-parallel topology, so as to identify the common data characteristics of each cluster and simultaneously perform hierarchical filtering processing on the raw data uploaded by each cluster to generate corresponding ordered data subsets.

[0135] The construction module is used to detect the amount of data to be processed and the task priority based on the ordered data subset, so as to construct the corresponding data statistics window and simultaneously build the collaborative communication link between the power plant monitoring terminals to construct the corresponding terminal data association matrix.

[0136] The deduction module is used to output the corresponding preliminary statistical results according to the data statistics window, and simultaneously, based on the terminal data association matrix, to reverse-deduce the preliminary statistical results according to the operating physical mechanism of the photovoltaic power station, so as to generate the corresponding deduction data sequence.

[0137] An integration module is used to calculate the degree of fit between the preliminary statistical results and the inferred data sequence, and simultaneously integrate the preliminary statistical results into a corresponding target statistical report when the degree of fit is detected to meet preset requirements.

[0138] Furthermore, the building module is specifically used for:

[0139] The short-term fluctuation correlation and long-term trend correlation in the ordered data subset are mined by the temporal attention mechanism to generate corresponding temporal effectiveness weights, and the amount of data to be processed is extracted from the ordered data subset according to the temporal effectiveness weights.

[0140] The amount of data to be processed is associated with the fault prediction risk level to generate the corresponding task priority. Simultaneously, based on the serial-parallel topology, the amount of data to be processed and the task priority are merged into a window construction unit.

[0141] Constrained by the bandwidth carrying capacity of the photovoltaic power station, the data sampling frequency and data buffer capacity of the window construction unit are dynamically matched, and the topology branch loss compensation coefficient is embedded synchronously to construct the data statistics window accordingly.

[0142] Furthermore, the building module is specifically used for:

[0143] Collect the actual temperature and voltage difference data of the photovoltaic modules corresponding to each of the power station monitoring terminals to calculate the corresponding transmission risk level, and simultaneously divide the power station monitoring terminals with similar transmission risk levels and adjacent physical locations into collaborative communication units.

[0144] Receive target monitoring data sent by each of the aforementioned collaborative communication units, and synchronously verify the integrity of the target monitoring data through hash value verification.

[0145] If the target monitoring data is found to be complete, the corresponding core association features are extracted, and each core association feature is simultaneously processed into a matrix using a sparse matrix reconstruction algorithm to generate the corresponding terminal data association matrix.

[0146] Furthermore, the inference module is specifically used for:

[0147] The operating physical mechanism of the photovoltaic power station is transformed into hard boundary constraints for reverse deduction. At the same time, topological features strongly correlated with the hard boundary constraints are extracted from the terminal data association matrix to construct a two-dimensional deduction baseline with mechanism constraints and topological features.

[0148] Using the abnormal data nodes in the preliminary statistical results as the starting point for tracing the source, and combining the dual-dimensional inference baseline, the historical operation data and interaction relationships corresponding to the starting point are traced through the terminal data association matrix. Combined with the operation physical mechanism of the photovoltaic power station, the source cause and propagation path of the data anomaly are deduced in reverse, and the corresponding tracing and inference path is generated.

[0149] The source tracing and deduction path is dynamically calibrated to output the corresponding deduction data sequence.

[0150] Furthermore, the inference module is specifically used for:

[0151] The operating physical mechanism of the photovoltaic power station is transformed into a corresponding inverse mapping function, and the topological correlation of the terminal data correlation matrix is ​​combined simultaneously to select effective nodes in the source tracing and deduction path.

[0152] Extract the temporal operation characteristics of the effective nodes, and simultaneously match the temporal anchoring nodes corresponding to the effective nodes in the terminal data association matrix according to the temporal operation characteristics;

[0153] The timing and numerical deviations of the effective nodes are dynamically corrected using the timing reference data of the timing anchor nodes to generate corresponding target nodes. Simultaneously, the target nodes are integrated and processed to generate the inference data sequence.

[0154] Furthermore, the integration module is specifically used for:

[0155] The preliminary statistical results are traced back to the ordered subset of data, and the inferred data sequence is simultaneously traced back to the terminal data association matrix to construct a dual-end tracing link.

[0156] By comparing the core data features in the dual-end traceability link, cross-verification is performed simultaneously to eliminate abnormal deviation data without traceability support, and the corresponding initial fit is calculated.

[0157] The initial fit is dynamically corrected to generate the corresponding fit.

[0158] Furthermore, the integration module is specifically used for:

[0159] Based on the terminal data association matrix, the implicit association factors between the various power station monitoring terminals are mined, and the corresponding coupling correction vector is created by combining the core parameters of the operating physical mechanism of the photovoltaic power station.

[0160] A correction benchmark is generated based on the characteristics of the same source data to match the initial fit, and the corresponding correction coefficient is matched simultaneously based on the magnitude of the correction benchmark.

[0161] Based on the correction coefficient and the coupling correction vector, a corresponding target correction vector is generated, and the target correction vector is simultaneously fused with the initial fit to generate the corresponding fit.

[0162] The fourth embodiment of the present invention provides a computer, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the photovoltaic power plant operation data statistics method as described above.

[0163] The fifth embodiment of the present invention provides a readable storage medium having a computer program stored thereon, wherein when the program is executed by a processor, it implements the photovoltaic power plant operation data statistics method as described above.

[0164] In summary, the photovoltaic power plant operation data statistics method and system provided in the above embodiments of the present invention can accurately and completely collect various types of data, thereby improving the efficiency of data statistics.

[0165] It should be noted that the above modules can be functional modules or program modules, and can be implemented through software or hardware. For modules implemented through hardware, the above modules can reside in the same processor; or the above modules can be located in different processors in any combination.

[0166] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-including system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0167] More specific examples of computer-readable media (a non-exhaustive list) include: electrical connections (electronic devices) having one or more wires, portable computer disk drives (magnetic devices), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Furthermore, computer-readable media can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in computer memory.

[0168] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0169] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0170] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the appended claims.

Claims

1. A method for statistical analysis of photovoltaic power plant operation data, characterized in that, The method includes: The photovoltaic modules are divided into several module clusters according to their installation area and series-parallel topology to identify the common data characteristics of each cluster. The raw data uploaded by each cluster is simultaneously processed by hierarchical filtering to generate corresponding ordered data subsets. Based on the ordered data subset, the amount of data to be processed and the task priority are detected to construct a corresponding data statistics window. At the same time, a collaborative communication link between power plant monitoring terminals is established to construct a corresponding terminal data association matrix. Based on the preliminary statistical results output by the data statistics window, and simultaneously based on the terminal data association matrix, the preliminary statistical results are reverse-engineered according to the operating physical mechanism of the photovoltaic power station to generate the corresponding inference data sequence. The degree of fit between the preliminary statistical results and the inferred data sequence is calculated, and when the degree of fit is detected to meet the preset requirements, the preliminary statistical results are integrated into the corresponding target statistical report. The steps of establishing a collaborative communication link between power plant monitoring terminals to construct a corresponding terminal data association matrix include: Collect the actual temperature and voltage difference data of the photovoltaic modules corresponding to each of the power station monitoring terminals to calculate the corresponding transmission risk level, and simultaneously divide the power station monitoring terminals with similar transmission risk levels and adjacent physical locations into collaborative communication units. Receive target monitoring data sent by each of the aforementioned collaborative communication units, and synchronously verify the integrity of the target monitoring data through hash value verification. If the target monitoring data is detected to be complete, the corresponding core association features are extracted, and each core association feature is simultaneously processed into a matrix using a sparse matrix reconstruction algorithm to generate the terminal data association matrix. The step of generating a corresponding deduced data sequence by reverse-engineering the preliminary statistical results based on the terminal data association matrix according to the operating physical mechanism of the photovoltaic power station includes: The operating physical mechanism of the photovoltaic power station is transformed into hard boundary constraints for reverse deduction. At the same time, topological features strongly correlated with the hard boundary constraints are extracted from the terminal data association matrix to construct a two-dimensional deduction baseline with mechanism constraints and topological features. Using the abnormal data nodes in the preliminary statistical results as the starting point for tracing the source, and combining the dual-dimensional inference baseline, the historical operation data and interaction relationships corresponding to the starting point are traced through the terminal data association matrix. Combined with the operation physical mechanism of the photovoltaic power station, the source cause and propagation path of the data anomaly are deduced in reverse, and the corresponding tracing and inference path is generated. The source tracing and deduction path is dynamically calibrated to output the deduction data sequence accordingly; The step of dynamically calibrating the source tracing and deduction path to output the deduction data sequence includes: The operating physical mechanism of the photovoltaic power station is transformed into a corresponding inverse mapping function, and the topological correlation of the terminal data correlation matrix is ​​combined simultaneously to select effective nodes in the source tracing and deduction path. Extract the temporal operation characteristics of the effective nodes, and simultaneously match the temporal anchoring nodes corresponding to the effective nodes in the terminal data association matrix according to the temporal operation characteristics; The timing and numerical deviations of the effective nodes are dynamically corrected using the timing reference data of the timing anchor nodes to generate corresponding target nodes. Simultaneously, the target nodes are integrated and processed to generate the inference data sequence.

2. The photovoltaic power plant operation data statistical method according to claim 1, characterized in that, The step of detecting the amount of data to be processed and the task priority based on the ordered data subset to construct the corresponding data statistics window includes: The short-term fluctuation correlation and long-term trend correlation in the ordered data subset are mined by the temporal attention mechanism to generate corresponding temporal effectiveness weights, and the amount of data to be processed is extracted from the ordered data subset according to the temporal effectiveness weights. The amount of data to be processed is associated with the fault prediction risk level to generate the corresponding task priority. Simultaneously, based on the serial-parallel topology, the amount of data to be processed and the task priority are merged into a window construction unit. Constrained by the bandwidth carrying capacity of the photovoltaic power station, the data sampling frequency and data cache capacity of the window construction unit are dynamically matched, and the topology branch loss compensation coefficient is embedded synchronously to construct the data statistics window accordingly.

3. The photovoltaic power plant operation data statistical method according to claim 1, characterized in that, The step of calculating the degree of fit between the preliminary statistical results and the inferred data sequence includes: The preliminary statistical results are traced back to the ordered subset of data, and the inferred data sequence is simultaneously traced back to the terminal data association matrix to construct a dual-end tracing link. By comparing the core data features in the dual-end traceability link, cross-verification is performed simultaneously to eliminate abnormal deviation data without traceability support, and the corresponding initial fit is calculated. The initial fit is dynamically corrected to generate the corresponding fit.

4. The photovoltaic power plant operation data statistical method according to claim 3, characterized in that, The step of dynamically correcting the initial fit to generate the corresponding fit includes: Based on the terminal data association matrix, the implicit association factors between the various power station monitoring terminals are mined, and a corresponding coupling correction vector is created by combining the core parameters of the operating physical mechanism of the photovoltaic power station. A correction benchmark is generated based on the characteristics of the same source data to match the initial fit, and the corresponding correction coefficient is matched simultaneously based on the magnitude of the correction benchmark. Based on the correction coefficient and the coupling correction vector, a corresponding target correction vector is generated, and the target correction vector is simultaneously fused with the initial fit to generate the corresponding fit.

5. A photovoltaic power plant operation data statistics system, characterized in that, The system is used to implement the photovoltaic power plant operation data statistics method as described in any one of claims 1 to 4, the system comprising: The identification module is used to divide the photovoltaic modules into several module clusters according to their installation area and series-parallel topology, so as to identify the common data characteristics of each cluster and simultaneously perform hierarchical filtering processing on the raw data uploaded by each cluster to generate corresponding ordered data subsets. The construction module is used to detect the amount of data to be processed and the task priority based on the ordered data subset, so as to construct the corresponding data statistics window and simultaneously build the collaborative communication link between the power plant monitoring terminals to construct the corresponding terminal data association matrix. The deduction module is used to output the corresponding preliminary statistical results according to the data statistics window, and simultaneously, based on the terminal data association matrix, to reverse-deduce the preliminary statistical results according to the operating physical mechanism of the photovoltaic power station, so as to generate the corresponding deduction data sequence. An integration module is used to calculate the degree of fit between the preliminary statistical results and the inferred data sequence, and simultaneously integrate the preliminary statistical results into a corresponding target statistical report when the degree of fit is detected to meet preset requirements.

6. A computer comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the photovoltaic power plant operation data statistics method as described in any one of claims 1 to 4.

7. A readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements the photovoltaic power plant operation data statistics method as described in any one of claims 1 to 4.