Photovoltaic project survey data integration and analysis platform

By designing a photovoltaic project survey data integration and analysis platform, using deep fusion technology and deep learning models, the problems of data dispersion and isolation in the existing technology and the lack of adaptability of the model are solved, and the deep correlation analysis of photovoltaic project survey data and accurate power generation potential prediction are achieved.

CN120030307AActive Publication Date: 2025-05-23GUANGDONG KUNLUN DIGITAL INTELLIGENT SOURCE TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510156299.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-12
Publication Date
2025-05-23
Estimated Expiration
2045-02-12

AI Technical Summary

Technical Problem

The existing technology is difficult to deeply analyze photovoltaic project survey data from different sources, resulting in dispersed and isolated data and cannot fully support project decision-making. At the same time, the analysis model lacks adaptability and makes it difficult to accurately predict the potential of photovoltaic power generation.

Method used

A photovoltaic project survey data integration and analysis platform was designed, including a multi-source data acquisition module, a multi-source data fusion module, an adaptive model construction module and a data analysis and prediction module. Deep fusion technology and deep learning model are used to achieve deep fusion of data and adaptive adjustment of models through technical means such as semantic analysis, convolutional neural network and principal component analysis.

Benefits of technology

In-depth correlation analysis of survey data of different photovoltaic projects has been achieved, the data fusion and analysis accuracy have been improved, the photovoltaic power generation potential can be predicted more accurately, and the decision-making of each stage of the project is supported.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030307A_ABST
    Figure CN120030307A_ABST
Patent Text Reader

Abstract

The invention relates to a photovoltaic project survey data integration and analysis platform, which is characterized in that an artificial intelligence model is adopted to learn the relevance among semantic analysis, image features and data scores of survey data of different regions, and the relevance is used as a relevance mode of the survey data of the different regions; training a deep learning model between the survey data and the photovoltaic generating capacity by using historical data, and adaptively adjusting the deep learning model according to the association mode of the survey data of the photovoltaic project to be evaluated; according to the method, the problem that the power generation potential evaluation of the photovoltaic project by the constructed model is inaccurate due to the fact that the difference of association modes among survey data of different regions cannot be considered in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of photovoltaic power generation, and relates to a photovoltaic project survey data integration and analysis platform. Background Art

[0002] With the increasing global demand for clean energy, photovoltaic projects, as an important part of the renewable energy field, are booming. During the implementation of photovoltaic projects, the analysis of preliminary survey data can evaluate and predict the potential of photovoltaic power generation, which plays a decisive role in the planning, construction, and operation of the projects.

[0003] However, there are many shortcomings in current technical means. On the one hand, the data integration degree is poor. The sources of preliminary survey data for projects are extensive, covering multiple aspects such as meteorology, geographical information, and on-site measurement. Existing technologies are difficult to deeply analyze and correlate the data from these different sources (regions), resulting in the data being in a scattered and isolated state and unable to provide comprehensive support for project decision-making. On the other hand, the analysis model lacks adaptability. The environments where photovoltaic projects are located are complex and changeable, and there are huge differences in climate, terrain, geology, etc. among different regions. Most existing analysis models are constructed based on fixed parameters and are difficult to automatically adjust according to the unique environmental characteristics of the project, resulting in poor accuracy and reliability of the analysis results and unable to accurately guide the work in each stage of the project. Summary of the Invention

[0004] The present invention provides a photovoltaic project survey data integration and analysis platform, aiming to achieve in-depth data fusion analysis and model adaptive adjustment, and solve the problem of inaccurate prediction of photovoltaic power generation potential due to differences in the correlation patterns of survey data for different photovoltaic projects.

[0005] The object of the present invention can be achieved through the following technical solutions: The present application provides a photovoltaic project survey data integration and analysis platform, including a multi-source data acquisition module, a multi-source data fusion module, an adaptive model construction module, and a data analysis and prediction module. The multi-source data acquisition module, the multi-source data fusion module, the adaptive model construction module, and the data analysis and prediction module are communicatively connected, wherein: The multi-source data acquisition module is used to collect the survey data of the area where the photovoltaic project is located; The multi-source data fusion module is used to analyze and process the survey data by using deep fusion technology and identify the correlation pattern of the survey data; The adaptive model construction module is used to train a deep learning model between the survey data and the photovoltaic power generation amount by using historical data and adaptively adjust the deep learning model according to the correlation pattern; The data analysis and prediction module is used to input the survey data of the photovoltaic project to be evaluated into the deep learning model and output the predicted photovoltaic power generation amount.

[0006] Furthermore, the survey data includes meteorological data, geographic data, equipment data and aerial images.

[0007] Furthermore, the deep fusion technology includes: S1. Obtain samples of PV project survey datasets in different regions during historical periods; S2, use NLP technology to perform semantic analysis on historical survey data; S3, using a convolutional neural network model to extract image features of historical survey data; S4. Calculate the data scores of historical survey data using principal component analysis; S5. Use artificial intelligence models to learn the semantic analysis of historical survey data from different regions, the correlation between image features and data scores, as the correlation model after the fusion of survey data from different regions.

[0008] Furthermore, the use of NLP technology to perform semantic analysis on historical survey data includes the following steps: Data preprocessing: text cleaning, normalization and word segmentation of historical survey data; Part-of-speech tagging: Attach a part-of-speech tag to each processed word to clarify the grammatical relationship between words in the sentence; Named Entity Recognition: Extract named entities related to PV project survey from word segmentation; Syntactic analysis: Analyze the grammatical structure of a sentence through dependency syntactic analysis and analyze the dependency relationship between each participle to understand the semantics of the sentence; Semantic role labeling: Based on syntactic analysis, we deeply explore the semantic relationship between the predicate and other components in the sentence and clarify the semantic role of each component; Knowledge fusion and semantic understanding: Integrate the information obtained from the previous steps and deeply integrate it with the professional knowledge map in the photovoltaic field to convert text data into structured knowledge related to photovoltaic project survey and analysis.

[0009] Furthermore, the method of extracting image features of historical survey data using a convolutional neural network model comprises the following steps: Data preparation: resize the image data of historical survey data to a standard size, convert it to color or grayscale, and normalize the pixel values; Construct a convolutional neural network model: Define the network layers based on the task and data characteristics, including the convolution layer for extracting local features, the pooling layer for downsampling to reduce the amount of data, the fully connected layer for mapping features, and the activation function for enhancing nonlinear expression; Model training: Divide the preprocessed image data into training set, validation set and test set in proportion, set the learning rate, batch size and number of training rounds to train the parameters, use the training set to train the model, calculate the loss through forward propagation, and update the parameters through back propagation. During training, use the validation set to evaluate the performance and adjust the parameters, and finally use the test set to evaluate the generalization ability of the model; Feature extraction: After the model training meets the standards, the historical survey images are input into the model, and the selected feature layer output, i.e., image features, is obtained through forward propagation.

[0010] Furthermore, the method of calculating the data score of the historical survey data by using principal component analysis comprises the following steps: Data standardization: The Z-score standardization method was used to standardize the values ​​of each variable; Calculate the covariance matrix: After completing data standardization, calculate the covariance matrix for historical survey data; Solving eigenvalues ​​and eigenvectors: Solving eigenvalues ​​and eigenvectors of covariance matrix expansion; Principal component selection: select the principal components based on the solved eigenvalues ​​and eigenvectors; Calculate data scores: Calculate data scores using the selected principal components and standardized historical survey data.

[0011] Furthermore, the artificial intelligence model is used to learn the semantic parsing of historical survey data in different regions, the correlation between image features and data scores, including the following steps: Data integration and preparation: The semantic analysis results, image features and data scores of historical survey data in different regions are fully integrated and sorted by region; Model selection: Select an AI model based on the characteristics of survey data in different regions and the complexity of the association; Model training: Using semantic analysis, image features, and data scores from different regions as input, the model learns the respective association patterns between data from different regions.

[0012] Furthermore, the deep learning model is configured as a multi-layer perceptron, including the following construction steps: Determine the input layer: Determine the number of neurons in the input layer according to the survey data dimension, and use the original data as input to provide basic information for model learning; Design hidden layers: The number of hidden layers and neurons in each layer are determined through experiments and experience. The layers are connected by weight matrices, and the neurons are transformed nonlinearly through weighted summation and activation functions. Construct the output layer: Set a neuron in the output layer to receive the output of the last hidden layer, and obtain the predicted value through weighted summation and linear activation function; Initialize weights and biases: After the model is built, initialize the weight matrix and neuron bias of each layer, and the initialization uses random initialization weights; Select loss function and optimizer: Select mean square error loss function to measure the difference between predicted and actual power generation, and use stochastic gradient descent method to feedback and adjust weight bias to minimize the loss value.

[0013] Furthermore, the adaptive adjustment of the deep learning model according to the association pattern comprises the following steps: Correlation pattern analysis: Use visualization and statistical analysis methods to clarify the interaction between survey data and the degree of influence on power generation; Feature importance assessment: Based on the association pattern, the random forest model is used to determine the importance of different associated survey data feature combinations in predicting power generation, and to quantify the influence of each survey data feature; Model structure adjustment: Based on the feature importance evaluation results, the deep learning model structure is optimized, specifically by increasing the neurons or nodes of important features and reducing the parts related to unimportant features; Weight and parameter adjustment: Increase the adjustment range of weights and parameters of important features; Model validation and iteration: Use the validation data set to test the adjusted model and use evaluation indicators to judge the model prediction performance.

[0014] Beneficial effects of the present invention: (1) An artificial intelligence model is used to learn the semantic analysis, image features and correlation between data scores of survey data in different regions as the correlation pattern of survey data in different regions; historical data is used to train a deep learning model between survey data and photovoltaic power generation, and the deep learning model is adaptively adjusted according to the correlation pattern of survey data of the photovoltaic project to be evaluated; the present invention solves the problem that the prior art cannot take into account the differences in correlation patterns between survey data in different regions, resulting in inaccurate evaluation of the power generation potential of photovoltaic projects by the constructed model.

[0015] (2) The survey data includes meteorological data, geographic data, equipment data and aerial images. The identified survey data correlation patterns can accurately characterize the power generation differences in different regions, thereby improving the accuracy of photovoltaic project evaluation.

[0016] (3) The model is adaptively adjusted based on the impact of different survey data characteristics on photovoltaic power generation in different regions, thereby improving the accuracy of model prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to facilitate understanding by those skilled in the art, the present invention is further described below with reference to the accompanying drawings.

[0018] Figure 1This is a structural diagram of a photovoltaic project survey data integration and analysis platform in the present invention.

[0019] Figure 2 This is a flowchart of deep fusion technology analysis in one embodiment of the present invention.

[0020] Figure 3 This is a flowchart of adaptively adjusting a deep learning model according to an association pattern in one embodiment of the present invention. DETAILED DESCRIPTION

[0021] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the specific implementation mode, structure, characteristics and effects of the present invention are described in detail below in combination with the accompanying drawings and preferred embodiments.

[0022] See also Figure 1-Figure 3 The present application provides a photovoltaic project survey data integration and analysis platform, including a multi-source data acquisition module, a multi-source data fusion module, an adaptive model construction module and a data analysis and prediction module, wherein the multi-source data acquisition module, the multi-source data fusion module, the adaptive model construction module and the data analysis and prediction module are communicatively connected, wherein: The multi-source data acquisition module is used to collect survey data of the area where the photovoltaic project is located; Furthermore, the survey data includes meteorological data, geographic data, equipment data and aerial images.

[0023] In this embodiment, the meteorological data includes information such as light intensity, sunshine duration, temperature, wind speed and wind direction. Light intensity and sunshine duration directly determine the total amount of solar energy received and converted into electricity by photovoltaic modules, and are key factors affecting power generation; temperature changes affect the photovoltaic conversion efficiency of modules, and excessively high or low temperatures may reduce power generation; wind speed and wind direction are related to equipment heat dissipation and stability, and strong winds may put pressure on structures such as photovoltaic brackets.

[0024] Geographic data includes topography, altitude, longitude and latitude, etc. Topography affects the feasibility and layout of power station construction. Flat terrain is conducive to large-scale installation, while complex terrain requires consideration of slope, aspect and other aspects to optimize component arrangement. Altitude affects atmospheric transparency and light intensity. The higher the altitude, the greater the light intensity. Longitude and latitude determine the solar altitude angle and sunshine duration, which affects power generation.

[0025] Equipment data covers the parameters and operating status information of various equipment in the photovoltaic project, such as the type, power, and conversion efficiency of photovoltaic modules, the conversion efficiency of inverters, and the maximum power tracking accuracy, as well as the temperature, current, voltage and other operating status data of the equipment. These data directly determine the power generation capacity, power conversion efficiency and normal operation of the equipment, and are crucial to ensuring the stable operation of the power generation system. Aerial images provide a macroscopic perspective for photovoltaic projects and can intuitively display the overall situation of the site, including topography, surrounding environment, building distribution, etc. By analyzing aerial images, shadow-blocked areas can be quickly identified, such as the impact of surrounding buildings and trees on the power station, so as to optimize the planning layout; it can also assist in evaluating land use and determine areas suitable for installing components.

[0026] The multi-source data fusion module is used to analyze and process the survey data using deep fusion technology to identify the correlation pattern of the survey data; In this embodiment, the deep fusion technology is committed to comprehensively analyzing the survey data of photovoltaic projects and exploring the inherent connections between different types of data. In view of the differences in the data association patterns of photovoltaic projects in different regions, this technology uses multi-step operations to accurately extract the association patterns applicable to specific projects, laying a solid foundation for subsequent data processing and analysis. Including: S1. Obtain samples of survey data sets of photovoltaic projects in different regions in historical periods: This step is the starting point of deep fusion technology, and it is necessary to widely collect survey data of photovoltaic projects in different historical periods and diverse geographical locations. These data come from a wide range of sources, covering meteorological data (such as light intensity, temperature, precipitation, etc.), geographic data (including altitude, topography, etc.), equipment data (such as photovoltaic module specifications, inverter performance parameters) and aerial images. Comprehensive and accurate sample data is the cornerstone of subsequent analysis. Sample data from different regions can reflect the characteristics of photovoltaic projects in various places, such as data differences between high-altitude areas and low-altitude areas, and arid areas and humid areas. These differences are crucial to revealing the differences in correlation patterns of survey data in different places.

[0027] S2. Use NLP technology to perform semantic analysis on historical survey data: Among the numerous survey data, multi-source data may contain text information. For the weather forecast text in meteorological data, NLP technology can split the words such as "tomorrow there will be thunderstorms in some areas, and the wind speed is expected to be 5-7 levels" into "tomorrow", "some areas", "there will be", "thunderstorms", "expected", "wind speed", "5-7 levels" and other vocabulary units through word segmentation; part-of-speech tagging determines the part-of-speech of each word, such as "thunderstorms" is a noun and "5-7 levels" is a quantifier; named entities identify key information, such as the weather phenomenon "thunderstorms" and the wind speed value "5 - 7 levels"; semantic role tagging clarifies the semantic role of words, such as "5 - 7 levels" is the specific value of "wind speed". In the text related to geographic data, such as the description of topography in the project site selection report, NLP technology can extract key information such as terrain type and terrain undulation. For the text of equipment data, such as the equipment maintenance manual, NLP technology can parse out the equipment failure type, maintenance cycle and other content.

[0028] Furthermore, the use of NLP technology to perform semantic analysis on historical survey data includes the following steps: Data preprocessing: This is the first step in using NLP technology to perform semantic analysis on historical survey data. Since historical survey data comes from a wide range of sources and in various forms, the text therein is often mixed with various types of interference information. Text cleaning aims to remove noise such as garbled characters, special characters, and HTML tags in web page texts, so that the text returns to a pure state. Normalization focuses on unifying the format of key information such as numbers, dates, and times in the text, and restoring abbreviations to avoid understanding difficulties caused by format differences or brief expressions. Word segmentation is a key part of data preprocessing. Based on different language characteristics, professional tools are used to accurately segment continuous text into independent words or word units, allowing computers to process text at a finer granularity, laying the foundation for subsequent steps such as part-of-speech tagging and entity recognition.

[0029] Part-of-speech tagging: After completing data preprocessing, part-of-speech tagging attaches a part-of-speech label to each word to clarify its role in the grammatical system, such as noun, verb, adjective, etc. This process is crucial for a deep understanding of text structure and semantic relationships. Through part-of-speech tagging, we can gain insight into the grammatical associations between words in a sentence, such as determining the main components such as subject, predicate, and object. Taking the photovoltaic project text "Inverter continuously outputs stable current" as an example, "inverter" is tagged as a noun as the subject of the action; "output" is tagged as a verb, acting as a predicate to represent the core action; "current" is also a noun and is the object of the action. Part-of-speech tagging tools are mostly based on statistical models or deep learning models, which can learn the part-of-speech rules of words based on a large amount of text data, thereby accurately marking the part of speech.

[0030] Named entity recognition: This step focuses on accurately extracting entities with specific meanings from the text. These entities are closely related to the survey of photovoltaic projects, covering key information such as project location, equipment model, measurement time, meteorological data values, etc. Deep learning-based methods, such as the Bi-LSTM + CRF model, perform well in processing complex text structures and diverse entity types, and can effectively identify various entities, providing key information support for subsequent semantic analysis.

[0031] Syntactic analysis: Syntactic analysis aims to analyze the grammatical structure of a sentence and clarify the dependency relationship between various components, such as subject, predicate, object, attributive, adverbial, complement, etc. Through this analysis, the organizational structure and semantic logic of the sentence can be clarified, just like building a skeleton for the text. Dependency syntactic analysis and phrase structure analysis are commonly used methods. Dependency syntactic analysis can clearly show the dependency relationship between words, assist in understanding the semantics of the sentence, and lay a solid foundation for more in-depth semantic analysis.

[0032] Semantic role labeling: Based on syntactic analysis, this step deeply explores the semantic relationship between the predicate (mostly verbs) and other components in the sentence, and clarifies the semantic role of each component, such as agent, patient, time, place, etc. Semantic role labeling relies on deep learning models and combines part-of-speech tagging, named entity recognition and syntactic analysis results to accurately predict the semantic role of each component and help fully understand the semantic connotation of the text.

[0033] Knowledge fusion and semantic understanding: As the final step of semantic analysis, this link integrates the information obtained from the previous steps and deeply integrates it with the professional knowledge map in the photovoltaic field. For example, the identified equipment model is associated with the detailed parameters, performance characteristics, common faults and other information of the equipment model in the knowledge map; the project location is associated with the geographical features, meteorological conditions, policy environment and other knowledge of the area in the knowledge map. Through this fusion, text information is no longer isolated, but intertwined with domain knowledge, so as to dig out the deep semantics behind the text, such as the potential causal relationship between different factors, best practice recommendations under specific conditions, etc., and convert text data into structured knowledge with practical value for photovoltaic project survey analysis and decision-making.

[0034] S3. Use convolutional neural network model to extract image features of historical survey data: Convolutional neural network plays a key role in data such as aerial images that directly present the status of the project area. It automatically extracts key features in the image, such as terrain undulations, vegetation coverage, and distribution of surrounding buildings through operations such as convolution layers and pooling layers. Aerial images in different regions have their own characteristics due to different geographical environments. CNN can capture these differences. For example, aerial images in mountainous areas show complex terrain features, while images in plain areas are relatively flat and open. These features combined with other data help to discover the unique association between image features in different regions and other survey data, such as the relationship between mountain terrain features and light shielding and power generation.

[0035] Furthermore, the method of extracting image features of historical survey data using a convolutional neural network model comprises the following steps: Data preparation: First, we collect various historical survey images of photovoltaic projects, such as aerial photos and equipment pictures, and organize them according to rules. Then we pre-process the original images, resize them to standard sizes, convert them to color or grayscale, and normalize the pixel values ​​to prepare for the input of the convolutional neural network model, ensuring that the data retains key information and meets the model processing requirements.

[0036] Construct a convolutional neural network model: Select a suitable architecture based on the task and data characteristics, such as classic models such as VGG and ResNet. On this basis, define the network layer, including the convolution layer to extract local features, the pooling layer to downsample to reduce the amount of data, and the fully connected layer to map features, and use activation functions to enhance nonlinear expression to build a network structure that can effectively learn image features.

[0037] Model training: Divide the preprocessed image data into training set, validation set and test set in proportion. Set training parameters such as learning rate, batch size, number of training rounds, train the model with the training set, calculate the loss through forward propagation, and update the parameters through back propagation. During training, use the validation set to evaluate the performance and adjust the parameters to prevent overfitting. Finally, use the test set to evaluate the generalization ability of the model.

[0038] Feature extraction: After the model training reaches the standard, it is loaded and the appropriate layer is selected from each network layer of the model according to the needs. The shallow layer extracts low-level features and the deep layer extracts high-level features. The historical survey image is input into the model, and the selected feature layer output is obtained through forward propagation, that is, the image feature representation, which is used for subsequent photovoltaic project data analysis and mining of potential connections with other data.

[0039] S4. Use principal component analysis to calculate the data scores of historical survey data: Survey data contains multiple variables, and there may be complex relationships between variables. Principal component analysis aims to reduce the dimension of the data, remove redundant information, and extract key components. The original data is projected onto new orthogonal coordinate axes through linear transformation, and the principal components are determined based on the size of the variance. The data scores of each sample on the principal components are calculated, and these scores summarize the core characteristics of the samples. The relationship between data variables in different regions is different. For example, in some high-temperature areas, the relationship between temperature and power generation is close, while in high-altitude areas, the correlation between altitude and light intensity is more significant. PCA can help sort out these regional differences in data relationships and highlight the main feature combinations of data from different regions.

[0040] Furthermore, the method of calculating the data score of the historical survey data by using principal component analysis comprises the following steps: Data standardization: Historical survey data covers a variety of variables, with large differences in dimensions and numerical ranges. Direct analysis will cause some variables to dominate the results. To this end, a standardization method such as Z-score is used to convert the values ​​of each variable into standard normal distribution data with a mean of 0 and a standard deviation of 1, eliminating the impact of dimensions and ensuring that all variables are on a unified scale, laying a standardized data foundation for subsequent principal component analysis.

[0041] Calculate the covariance matrix: After completing data standardization, calculate the covariance matrix for historical survey data. This matrix displays the covariance values ​​between different variables through elements, comprehensively presents the linear correlation between various types of data such as meteorology, geography, and equipment, helps to clarify the degree of influence of variable combinations on the overall characteristics of the data, and provides core information for extracting the main components.

[0042] Solving eigenvalues ​​and eigenvectors: Solving the eigenvalues ​​and eigenvectors of the covariance matrix. The eigenvalues ​​reflect the amount of information of each principal component, and the eigenvectors determine the direction of the principal component. In historical survey data, the principal component corresponding to the larger eigenvalue contains the main information, which may be a linear combination of multiple original variables. This step extracts the key principal component direction and information content from the internal structure of the data.

[0043] Principal component selection: select the principal components based on the solved eigenvalues ​​and eigenvectors. Generally, they are sorted by eigenvalue size, and the first few principal components whose cumulative contribution rate reaches a specific threshold (such as 80% - 90%) are selected. The cumulative contribution rate represents the proportion of information contained in the selected principal component. In the PV project survey data, the appropriate selection of principal components can retain key information, reduce the dimension, and avoid excessive information loss without affecting subsequent analysis.

[0044] Calculate data scores: Finally, use the selected principal components and the standardized historical survey data to calculate the data scores. Perform matrix multiplication on the standardized data and the principal component eigenvectors to obtain the scores of each sample on each principal component. These scores reflect the position of the sample in the new principal component space and comprehensively reflect the characteristics of the sample under the original variable combination, providing a quantitative analysis basis for each link of the photovoltaic project.

[0045] S5. Use artificial intelligence models to learn the semantic analysis of historical survey data in different regions, the correlation between image features and data scores, as the correlation pattern after the fusion of survey data in different regions: With the help of artificial intelligence models (such as deep neural networks, random forests, etc.), with structured data after semantic analysis, image features extracted by CNN, and data scores calculated by PCA as input, the model learns the mapping relationship between these different types of features through a large amount of training. The input data features of different regions are different, and the model will learn the unique correlation patterns of various places. In the in-depth analysis of historical survey data of photovoltaic projects, there is a multi-dimensional and close correlation between semantic analysis, image features and data scores.

[0046] From the perspective of meteorology and geographical environment, semantically parsed meteorological text data, such as descriptions of information such as duration, intensity, frequency and magnitude of precipitation, are significantly correlated with geographical features obtained through aerial images. For example, if semantic parsing shows that a certain area has "sufficient sunlight all year round and a lot of snow in winter", and the aerial image shows that the area is a plain terrain with open terrain and no tall obstructions, this explains the reason for sufficient sunlight; and if the image shows that there are mountains around and the project is on the leeward slope, it may correspond to the description of snow in winter. The data score obtained through principal component analysis will comprehensively consider the climatic conditions conveyed by the semantics of the meteorological text and the geographical features reflected in the image, so as to generate a data score that reflects the suitability of the overall meteorological-geographical environment of the region for photovoltaic projects. If the score is high, it means that the meteorological and geographical conditions in this area are more favorable to photovoltaic projects, such as sufficient sunlight and suitable terrain are conducive to the efficient power generation of photovoltaic modules; on the contrary, a low score may indicate that there are some problems that affect the project, such as snow in winter may cause photovoltaic modules to be blocked by snow, affecting power generation efficiency.

[0047] In terms of equipment performance and environmental factors, after semantic analysis of texts such as equipment maintenance manuals and fault reports, the information contained in them, such as the type of equipment failure, frequency, and cause of failure (such as "the output power of photovoltaic modules has dropped due to long-term high temperature"), is closely linked to the environmental conditions surrounding the equipment presented in the image features. For example, thermal imaging images may show that the surface temperature of photovoltaic modules is too high, which echoes the high temperature mentioned in the text that causes failures. The data score integrates the text semantic information related to equipment performance and the environmental characteristics reflected by the image, such as combining equipment failure conditions with factors such as ambient temperature, humidity, and wind and sand to derive a data score that reflects the operating stability of the equipment in the current environment. A low data score means that the equipment may frequently fail in the current environment, and it is necessary to further optimize the equipment cooling system, strengthen protective measures, or replace equipment models that are more suitable for the environment to ensure the stable operation of the photovoltaic project.

[0048] In the time series, the semantically parsed text data will contain information such as weather forecasts at different time points, equipment status updates, and project progress records; image features will also show dynamic changes over time, such as the growth and withering of vegetation in different seasons affecting the light shielding of photovoltaic components, and the changes in the appearance wear and aging of equipment during long-term use; data scores will also change due to time factors. Through in-depth analysis of the correlation of these data in time series, some long-term development trends and internal laws can be excavated. For example, as time goes by, signs of equipment aging can be found from the semantic parsing of equipment maintenance records, such as the gradual increase in failure frequency and the increase in maintenance complexity; at the same time, the appearance of the equipment in the image may show obvious aging characteristics such as wear and discoloration; the data score will also gradually decrease accordingly, comprehensively reflecting the deterioration of the overall operating condition of the equipment. This kind of correlation analysis in time series is of great significance for predicting the development trend of future photovoltaic projects. By learning the correlation patterns of historical data in the time dimension, the artificial intelligence model can accurately predict future changes in equipment performance, power generation fluctuations, and the probability of failure, thereby providing a scientific basis for the project's operation and maintenance decisions, such as formulating equipment replacement plans in advance and optimizing maintenance cycles; and providing strong support for resource allocation, such as reasonably arranging human and material resources to ensure that photovoltaic projects always maintain efficient and stable operation.

[0049] However, this correlation varies to a certain extent for different regions. For example, compared with low-altitude areas, the positive correlation between light and power generation in high-altitude areas is higher, and the weight of temperature on equipment performance is greater. Therefore, this embodiment uses an artificial intelligence model to learn the correlation between semantic analysis, image features and data scores of survey data in different regions to determine the correlation of survey data in different regions.

[0050] In the entire PV project survey data processing process, the previous steps respectively carried out multi-dimensional analysis of historical survey data, including obtaining samples, NLP semantic analysis, convolutional neural network extraction of image features, and principal component analysis to calculate data scores. Step S5 will focus on using artificial intelligence models to specifically learn the correlation between the above-mentioned types of data in different regions, so as to obtain survey data correlation patterns suitable for different regions, which makes the analysis results more region-specific. Including: Data integration and preparation: The semantic analysis results, image features, and data scores of the historical survey data obtained in the previous steps in different regions are fully integrated. The data for each region contains information extracted from the text, key features extracted from the image, and data scores obtained through comprehensive analysis. For example, the meteorological text in region A has unique climate characteristics after semantic analysis, and its image presents specific geographical features, as well as data scores based on various factors; region B also has corresponding different data. These data are sorted by region to form a data set that can be processed by the artificial intelligence model.

[0051] Choose the right AI model: Given the complexity of data characteristics and associations in different regions, it is necessary to select AI models that can effectively capture these differences. For situations where data associations are complex and have nonlinear relationships, deep neural network models such as multi-layer perceptrons (MLPs), recurrent neural networks (RNNs) and their variants, long short-term memory networks (LSTMs), and gated recurrent units (GRUs) may be more appropriate. For example, if there are complex and dynamically changing relationships between meteorological, geographical, and equipment data in different regions, LSTM or GRU can better handle such time series or sequential data associations. If the relationship between data is relatively simple, traditional machine learning models such as decision trees and random forests can also play a role. They can find out the association patterns between data in different regions by analyzing data features.

[0052] Model training: Use integrated data sets from different regions to train the selected artificial intelligence model. During the training process, the model uses semantic analysis, image features, and data scores from different regions as input to learn the unique association patterns between data from different regions. The model minimizes the error between the predicted association pattern and the actual association pattern by continuously adjusting its own parameters. For example, for a certain region, the model predicts the correlation between the light intensity, terrain features, and power generation in the region, compares it with the correlation in the actual data, and updates the parameters through the back-propagation algorithm to make the prediction results closer to the actual situation. The training process will be repeated on data from different regions, allowing the model to gradually grasp the association rules of data from different regions.

[0053] Associated Pattern Learning and Output: As the training progresses, the AI model gradually learns the associations between various features of historical survey data from different regions. These learned associations are the associated patterns of survey data from different regions. For example, the model may find that the correlation between light intensity and power generation in high-altitude areas is closer than that in low-altitude areas, or there are specific differences in the impact of humidity on equipment performance between coastal areas and inland areas. These associated patterns exist in the form of model parameters and internal representations and can be used to analyze and predict new survey data from the same region, providing targeted decision-making basis for the planning, construction, operation and maintenance of photovoltaic projects in different regions.

[0054] The adaptive model construction module is used to train a deep learning model between survey data and photovoltaic power generation using historical data, and adaptively adjust the deep learning model according to the associated pattern; In this embodiment, the adaptive model construction module aims to accurately construct and dynamically optimize the relationship model between survey data and photovoltaic power generation, focusing on training the deep learning model using historical data. By widely collecting various historical data related to photovoltaic projects, including meteorological survey data such as light intensity, temperature, humidity, etc.; geographical survey data such as altitude, topography, etc.; and equipment operation survey data such as component power, inverter efficiency, etc., while correspondingly recording the photovoltaic power generation data. After collection, preprocess these data, clean out outliers and noise data, and unify the scale of data with different dimensions through standardization processing to lay a solid foundation for subsequent model training. Based on data characteristics and project requirements, carefully select a suitable deep learning model architecture, such as multi-layer perceptron, recurrent neural network and its variants, etc., and build the model, determining key parameters such as the number of layers and neurons. Subsequently, use the preprocessed data to train the model, take the survey data as input and the power generation as the output label, and continuously adjust the model parameters with the help of optimization algorithms to gradually reduce the error between the predicted power generation and the actual power generation, thus constructing a preliminary prediction model.

[0055] Secondly, the module will adaptively adjust the deep learning model based on the association pattern. The association pattern here is derived from the multi-dimensional analysis of historical survey data, reflecting the intrinsic connection between various features of survey data in different regions. After deeply analyzing these association patterns and understanding the data relationship in different regions and how it affects power generation, the adaptive model building module will make targeted adjustments to the trained model. If it is found that the terrain factors in a certain area have a significant impact on power generation, and the initial model does not take it into account, the input layer or hidden layer related to the terrain characteristics will be added; or according to the importance of each factor in the association pattern, the feature weights in the model will be adjusted, and even the model structure will be adjusted according to the complexity of the association pattern. In this way, the model can more accurately fit the actual situation in different regions, significantly improve the prediction accuracy and adaptability of photovoltaic power generation, and provide strong support for the efficient planning, operation and maintenance, and decision-making of photovoltaic projects.

[0056] Furthermore, the deep learning model is configured as a multi-layer perceptron, including the following construction steps: Determine the input layer: The input layer of the multi-layer perceptron (MLP) needs to be built closely around the survey data of the photovoltaic project. The collection covers meteorological data (such as light intensity, temperature, humidity, wind speed, etc.), geographic data (such as altitude, slope, longitude and latitude, etc.), and equipment-related data (such as the model and power of photovoltaic modules, inverter efficiency, etc.). These data comprehensively reflect the various factors that affect photovoltaic power generation. According to the dimension of the data, determine the number of neurons in the input layer. For example, if a total of 10 different types of survey data are collected, then the input layer is set to 10 neurons, each neuron corresponds to a data feature, and these raw data are used as the input of the model to provide basic information for subsequent learning and prediction.

[0057] Designing hidden layers: The hidden layer is the core part of the MLP and is responsible for learning complex patterns in the data. It is critical to determine the number of hidden layers and the number of neurons in each layer. Generally, it is selected through experiments and experience, and 1-3 hidden layers are usually tried first. The number of neurons in each layer can be adjusted according to the number of neurons in the input layer and the complexity of the data. For example, if the data is more complex, more neurons may be set in the first hidden layer than in the input layer, such as 15-20, to increase the learning ability of the model. The number of neurons in subsequent hidden layers can be gradually reduced, such as 10-15 in the second layer and 5-10 in the third layer. The hidden layers are connected by weight matrices. Each neuron receives the output of the neurons in the previous layer by weighted summation, and is nonlinearly transformed by an activation function (such as the ReLU function), so that the model can learn the nonlinear relationship in the data and mine the complex relationship between survey data and power generation.

[0058] Construction of the output layer: The construction of the output layer is relatively simple. Since the goal is to predict the photovoltaic power generation, only one neuron needs to be set in the output layer. This neuron receives the output of the last hidden layer and obtains the final prediction result through weighted summation and an activation function (a linear activation function because the power generation is a continuous value). This predicted value is the photovoltaic power generation predicted by the model based on the input survey data. The weights between the output layer and the last hidden layer are also optimized through training to make the predicted power generation as close as possible to the actual power generation, thereby achieving the goal of accurately predicting photovoltaic power generation using the MLP model.

[0059] Initialization of weights and biases: After the model is constructed, it is necessary to initialize the weight matrices between layers and the biases of neurons. The way of weight initialization affects the training effect and convergence speed of the model. Common initialization methods include random initialization, such as using Gaussian distribution or uniform distribution to randomly generate weight values, ensuring that the weights take values within a certain range to avoid difficulties in training caused by overly large or small values. Biases are usually initialized to small constants, such as 0 or values close to 0. Reasonable initialization of weights and biases enables the model to start learning in a better state at the initial stage of training, laying a foundation for continuously optimizing weights and biases through the backpropagation algorithm in the subsequent process, thereby improving the prediction accuracy of the model.

[0060] Selection of loss function and optimizer: The loss function is used to measure the difference between the model's predicted value and the actual value. For the MLP model predicting photovoltaic power generation, the mean squared error (MSE) loss function is a commonly used choice. It calculates the average of the squares of the differences between the predicted power generation and the actual power generation, intuitively reflecting the accuracy of the model's prediction. The optimizer is responsible for adjusting the weights and biases of the model according to the feedback of the loss function to minimize the value of the loss function. Stochastic gradient descent (SGD) and its variants (such as Adagrad, Adadelta, Adam, etc.) are common optimizers. For example, the Adam optimizer combines the advantages of Adagrad and Adadelta, can adaptively adjust the learning rate, performs well during training, enables the model to converge to the optimal solution faster, and thus improves the prediction performance of the model.

[0061] Furthermore, the adaptive adjustment of the deep learning model according to the association pattern includes the following steps: Correlation pattern analysis: In-depth study of the correlation patterns between semantic analysis, image features and data scores in historical survey data from different regions. For example, analyze which factors have a significant impact on power generation in specific regions, such as the close connection between light intensity and power generation in high-altitude areas, or the impact of humidity and equipment performance on power generation in coastal areas. Through visualization tools (such as drawing heat maps to show the correlation between factors), statistical analysis (calculating correlation coefficients) and other methods, the interaction relationship between factors and the degree of impact on power generation are clearly presented, providing a clear basis for model adjustment.

[0062] Feature importance assessment: Determine the importance of different features in predicting power generation based on the association pattern. High importance is given to features that appear frequently in the association pattern and have a great impact on power generation. For example, in areas with abundant light resources, the light intensity feature is of high importance; while in areas with complex terrain, the importance of terrain-related features is more prominent. Use feature selection algorithms (such as recursive feature elimination) or feature importance assessment methods based on machine learning (such as feature importance ranking of random forests) to quantify the importance of each feature so that the model can be adjusted in a targeted manner later. In the data analysis of photovoltaic projects, the association pattern reflects the intrinsic connection between semantic parsing, image features and data scores in historical survey data. Random forest is an integrated learning model composed of multiple decision trees. The random forest model is used to determine the importance of different associated survey data feature combinations for predicting power generation based on the association pattern because it can process high-dimensional data, automatically consider the interaction between features, and has good stability and generalization ability. By training the random forest model, the influence of each survey data feature on power generation prediction can be quantified, so as to find the key features and their associated combinations that have a greater impact on power generation.

[0063] Model structure adjustment: According to the results of feature importance assessment, the structure of the multilayer perceptron (MLP) is adjusted. If a key feature is not fully reflected in the original model, the relevant input layer neurons or hidden layer nodes can be added. For example, if it is found that vegetation coverage in a certain area has a significant impact on power generation, and the original model does not consider this feature, vegetation coverage-related neurons can be added to the input layer, and nodes can be appropriately added to the hidden layer to enhance the learning ability of this feature. Conversely, for neurons or connections corresponding to features of low importance, it can be considered to be reduced or deleted to simplify the model structure and avoid overfitting.

[0064] Weight and parameter adjustment: In addition to structural adjustments, the weights and parameters of the model need to be optimized according to the association pattern. For the weights corresponding to the features with high importance, a larger adjustment range is given during the training process so that it can more accurately reflect the relationship between the feature and the power generation. For example, by adjusting the learning rate, a larger learning rate is used for the weights related to important features to accelerate its convergence. At the same time, re-examine the hyperparameters of the model (such as the number of hidden layers, neuron activation function, etc.), and make appropriate adjustments based on the characteristics of the association pattern and the performance of the model on the validation set to improve the model's adaptability to data from different regions and the accuracy of prediction.

[0065] Model verification and iteration: After completing the above adjustments, use the validation data set to verify the model. Run the adjusted model on the validation set, compare the predicted power generation with the actual power generation, and evaluate the model performance by calculating indicators such as mean square error (MSE) and mean absolute error (MAE). If the model performance does not meet expectations, repeat the above steps, further analyze the correlation pattern, adjust the model structure and parameters, and continue to iterate and optimize until the model shows good adaptability and prediction accuracy on the validation set, and can accurately reflect the relationship between survey data and power generation in different regions.

[0066] The above description is only a preferred embodiment of the present invention and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment as above, it is not used to limit the present invention. Any technical personnel in this field can make some changes or modify the technical contents disclosed above into equivalent embodiments without departing from the scope of the technical solution of the present invention. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention are still within the scope of the technical solution of the present invention.

Claims

1. A photovoltaic project survey data integration and analysis platform, characterized by: It includes a multi-source data acquisition module, a multi-source data fusion module, an adaptive model construction module and a data analysis and prediction module, which are communicatively connected to each other, wherein: The multi-source data acquisition module is used to collect survey data of the area where the photovoltaic project is located; The multi-source data fusion module is used to analyze and process the survey data using deep fusion technology to identify the correlation pattern of the survey data; The adaptive model building module is used to train a deep learning model between survey data and photovoltaic power generation using historical data, and adaptively adjust the deep learning model according to the association pattern; The data analysis and prediction module is used to input the survey data of the photovoltaic project to be evaluated into the deep learning model and output the predicted photovoltaic power generation.

2. A photovoltaic project survey data integration and analysis platform according to claim 1, characterized in that: The survey data includes meteorological data, geographic data, equipment data and aerial images.

3. The photovoltaic project survey data integration and analysis platform according to claim 1, characterized in that: The deep fusion technology includes: S1. Obtain samples of PV project survey datasets in different regions during historical periods; S2, use NLP technology to perform semantic analysis on historical survey data; S3, using a convolutional neural network model to extract image features of historical survey data; S4. Calculate the data scores of historical survey data using principal component analysis; S5. Use artificial intelligence models to learn the semantic analysis of historical survey data from different regions, the correlation between image features and data scores, as the correlation model after the fusion of survey data from different regions.

4. A photovoltaic project survey data integration and analysis platform according to claim 3, characterized in that: The method of using NLP technology to perform semantic analysis on historical survey data includes the following steps: Data preprocessing: text cleaning, normalization and word segmentation of historical survey data; Part-of-speech tagging: Attach a part-of-speech tag to each processed word to clarify the grammatical relationship between words in the sentence; Named Entity Recognition: Extract named entities related to PV project survey from word segmentation; Syntactic analysis: Analyze the grammatical structure of a sentence through dependency syntactic analysis and analyze the dependency relationship between each participle to understand the semantics of the sentence; Semantic role labeling: Based on syntactic analysis, we deeply explore the semantic relationship between the predicate and other components in the sentence and clarify the semantic role of each component; Knowledge fusion and semantic understanding: Integrate the information obtained from the previous steps and deeply integrate it with the professional knowledge map in the photovoltaic field to convert text data into structured knowledge related to photovoltaic project survey and analysis.

5. The photovoltaic project survey data integration and analysis platform according to claim 3, characterized in that: The method of extracting image features of historical survey data using a convolutional neural network model comprises the following steps: Data preparation: resize the image data of historical survey data to a standard size, convert it to color or grayscale, and normalize the pixel values; Construct a convolutional neural network model: Define the network layers based on the task and data characteristics, including the convolution layer for extracting local features, the pooling layer for downsampling to reduce the amount of data, the fully connected layer for mapping features, and the activation function for enhancing nonlinear expression; Model training: Divide the preprocessed image data into training set, validation set and test set in proportion, set the learning rate, batch size and number of training rounds to train the parameters, use the training set to train the model, calculate the loss through forward propagation, and update the parameters through back propagation. During training, use the validation set to evaluate the performance and adjust the parameters, and finally use the test set to evaluate the generalization ability of the model; Feature extraction: After the model training meets the standards, the historical survey images are input into the model, and the selected feature layer output, i.e., image features, is obtained through forward propagation.

6. The photovoltaic project survey data integration and analysis platform according to claim 3, characterized in that: The method of calculating the data score of the historical survey data by principal component analysis comprises the following steps: Data standardization: The Z-score standardization method was used to standardize the values ​​of each variable; Calculate the covariance matrix: After completing data standardization, calculate the covariance matrix for historical survey data; Solving eigenvalues ​​and eigenvectors: Solving eigenvalues ​​and eigenvectors of covariance matrix expansion; Principal component selection: select the principal components based on the solved eigenvalues ​​and eigenvectors; Calculate data scores: Calculate data scores using the selected principal components and standardized historical survey data.

7. The photovoltaic project survey data integration and analysis platform according to claim 3, characterized in that: The method of using an artificial intelligence model to learn the semantic parsing of historical survey data in different regions, the correlation between image features and data scores includes the following steps: Data integration and preparation: The semantic analysis results, image features and data scores of historical survey data in different regions are fully integrated and sorted by region; Model selection: Select an AI model based on the characteristics of survey data in different regions and the complexity of the association; Model training: Using semantic analysis, image features, and data scores from different regions as input, the model learns the respective association patterns between data from different regions.

8. The photovoltaic project survey data integration and analysis platform according to claim 1, characterized in that: The deep learning model is configured as a multi-layer perceptron and includes the following construction steps: Determine the input layer: Determine the number of neurons in the input layer according to the survey data dimension, and use the original data as input to provide basic information for model learning; Design hidden layers: The number of hidden layers and neurons in each layer are determined through experiments and experience. The layers are connected by weight matrices, and the neurons are transformed nonlinearly through weighted summation and activation functions. Construct the output layer: Set a neuron in the output layer to receive the output of the last hidden layer, and obtain the predicted value through weighted summation and linear activation function; Initialize weights and biases: After the model is built, initialize the weight matrix and neuron bias of each layer, and the initialization uses random initialization weights; Select loss function and optimizer: Select mean square error loss function to measure the difference between predicted and actual power generation, and use stochastic gradient descent method to feedback and adjust weight bias to minimize the loss value.

9. The photovoltaic project survey data integration and analysis platform according to claim 1, characterized in that: The step of adaptively adjusting the deep learning model according to the association pattern comprises the following steps: Correlation pattern analysis: Use visualization and statistical analysis methods to clarify the interaction between survey data and the degree of influence on power generation; Feature importance assessment: Based on the association pattern, the random forest model is used to determine the importance of different associated survey data feature combinations in predicting power generation, and to quantify the influence of each survey data feature; Model structure adjustment: Based on the feature importance evaluation results, the deep learning model structure is optimized, specifically by increasing the neurons or nodes of important features and reducing the parts related to unimportant features; Weight and parameter adjustment: Increase the adjustment range of weights and parameters of important features; Model validation and iteration: Use the validation data set to test the adjusted model and use evaluation indicators to judge the model prediction performance.

Citation Information

Patent Citations

  • Regional photovoltaic power generation power prediction method and system based on multi-source data

    CN115759453A

  • Photovoltaic power prediction method based on multi-scale space-time diagram attention convolutional network

    CN117154704A

  • System for managing renewable energy generator

    KR102498535B1

  • Method for training power generation amount prediction model of photovoltaic power station, power generation amount prediction method and device of photovoltaic power station, training system, prediction system and storage medium

    WO2020228568A1

  • Training optimization method for foreign exchange time series prediction

    WO2021082809A1