A method and system for urban photovoltaic potential assessment based on multi-modal deep learning
By employing multimodal deep learning methods, this study addresses the issues of insufficient information fusion and the disconnect between assessment results and engineering applications in the assessment of urban building rooftop photovoltaic potential. It achieves efficient and accurate multi-level potential assessment, supporting urban energy planning and distributed photovoltaic layout.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA CONSTR THIRD ENG BUREAU GRP CO LTD
- Filing Date
- 2026-05-28
- Publication Date
- 2026-06-30
AI Technical Summary
Existing technologies lack a multi-source information fusion mechanism in assessing the potential of rooftop photovoltaics in urban buildings, making it difficult to accurately characterize the development conditions of rooftop photovoltaics in complex urban environments, to balance the speed and accuracy of assessments at the urban scale, and to develop a multi-level potential assessment system oriented towards engineering applications.
A multimodal deep learning-based approach is adopted to acquire and preprocess multimodal data, and construct a multimodal deep learning model, including a roof identification module and a multimodal fusion module. Through multi-level potential assessment, radiation potential, installation potential, power generation potential, carbon emission reduction potential and economic potential are calculated to form multi-level potential assessment results.
It improves the accuracy of roof identification and the rationality of assessment results, enhances the batch processing capability at the urban scale, and the output results can directly serve urban energy planning and engineering applications, providing spatial differentiation and decision-making references.
Smart Images

Figure CN122311641A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of deep learning and renewable energy assessment technology, and in particular to a method and system for assessing urban photovoltaic potential based on multimodal deep learning. Background Technology
[0002] With the development of high-resolution remote sensing imagery, geographic information systems, and deep learning technologies, automatically identifying urban building rooftops using satellite imagery and assessing their potential for rooftop photovoltaic development has become an important technical approach in current research. However, considering the actual needs of assessing the potential of urban building rooftop photovoltaics, existing technologies still have the following three main shortcomings:
[0003] 1. Existing technologies lack a multi-source information fusion mechanism, making it difficult to accurately characterize the conditions for rooftop photovoltaic development.
[0004] Many existing methods for assessing the potential of urban rooftop photovoltaics rely primarily on single remote sensing images, single GIS layers, or single attribute data for analysis. They focus on identifying rooftop boundaries and estimating geometric areas, but lack a unified approach to acquiring, aligning, integrating, and collaboratively representing the diverse information that influences the conditions for rooftop photovoltaic development.
[0005] In reality, the potential of rooftop photovoltaics is not only related to roof area, but also closely related to factors such as building age, building function, roof type, roof slope, building height, surrounding shading, and local microclimate conditions. It may even be indirectly affected by residents' socioeconomic factors. Without unified integration of this multi-source information, it is difficult to accurately characterize the conditions for rooftop photovoltaic development in complex urban environments, which can easily lead to assessment results that are too rough and lack applicability.
[0006] 2. Existing technologies struggle to balance the speed and accuracy of urban-scale assessments.
[0007] Existing methods for assessing the potential of urban building rooftop photovoltaics can be broadly categorized into two types: one relies primarily on single remote sensing images or single modal data to quickly identify and extract roof outlines, roof areas, or usable rooftops; the other, based on identification, further combines radiation analysis, shading analysis, environmental correction, or simulation calculations to conduct a more detailed assessment of rooftop photovoltaic development potential. The former boasts high processing speed at the urban scale, making it suitable for large-scale, batch applications. However, due to relatively limited input information and analytical dimensions, it often yields only coarse assessment results, failing to accurately reflect the true development conditions of rooftops in complex built-up environments. While the latter can improve assessment accuracy to some extent, it typically requires more computational steps and subsequent analysis processes, resulting in higher overall computational costs and longer processing cycles, making it difficult to meet the practical needs of rapid assessment at the urban scale.
[0008] Therefore, existing technologies generally present a contradiction: when using only a single modality or simplified method, the processing speed is fast but the accuracy of the results is limited; while after introducing more analytical factors and correction processes, the evaluation results are more refined, but the overall efficiency decreases significantly, making it difficult to simultaneously balance evaluation speed and result accuracy at the city scale.
[0009] 3. Existing technologies lack a multi-level potential assessment system for engineering applications.
[0010] From a practical application perspective, assessing the potential of urban building rooftop photovoltaics requires answering not only questions like "What is the roof area?" or "How much solar radiation can be received?", but also "Which rooftops are suitable for installation?", "What is the maximum installation scale?", "What is the theoretical power generation?", "How much carbon emission can be reduced?", and "Is it economically feasible?". However, current technologies often only provide results at the level of rooftop identification, usable area estimation, or a single theoretical power generation figure, lacking a systematic assessment chain encompassing radiation resources, installation scale, technical power generation capabilities, carbon emission reduction benefits, and economic returns. This limits the direct application value of the results, while possessing research significance, in urban energy planning, distributed photovoltaic layout, project selection, and investment decisions. Summary of the Invention
[0011] This invention proposes a method and system for assessing urban photovoltaic potential based on multimodal deep learning, aiming to solve the technical problems existing in the current urban building rooftop photovoltaic potential assessment methods, such as insufficient integration of multi-factor information, low computational efficiency at the urban scale, and disconnect between assessment results and engineering applications.
[0012] In a first aspect, the present invention provides a method for assessing urban photovoltaic potential based on multimodal deep learning, comprising:
[0013] S1. Acquire multimodal data of the study area. The multimodal data includes at least satellite remote sensing images, microclimate images, and tabular data. Preprocess the multimodal data to construct a unified multimodal input dataset.
[0014] S2, Construct a multimodal deep learning model, the multimodal deep learning model including a roof recognition module for identifying building roofs from the satellite remote sensing image, and a multimodal fusion module for fusing the multimodal data; Train the multimodal deep learning model using the multimodal input dataset, the trained model is used to output a roof mask and at least one photovoltaic potential correction parameter;
[0015] S3. Based on the roof mask and the photovoltaic potential correction parameters, according to the preset hierarchical logic, the radiation potential representing solar energy resources, the installation potential representing the area where photovoltaic modules can be arranged, the technical power generation potential representing the theoretical power generation capacity, the carbon emission reduction potential representing environmental benefits, and the economic potential representing economic benefits are calculated in sequence to form a multi-level potential assessment result.
[0016] Furthermore, S1 specifically includes:
[0017] The satellite remote sensing images and microclimate images are subjected to coordinate transformation, resampling, cropping and normalization to obtain image modal data;
[0018] Missing value cleanup, category coding, and numerical standardization are performed on the table data to obtain table modal data, which includes building attribute data, spatial morphology data, and engineering parameter data;
[0019] The image modal data and the tabular modal data are aligned using spatial units or building objects, so that each roof sample to be evaluated corresponds to a set of multimodal data. Represented as:
[0020] ;
[0021] in, For the first Satellite images of a sample, For the first A collection of microclimate images of individual samples. For the first The table feature vector of each sample.
[0022] Furthermore, the roof identification module in S2 is a semantic segmentation network used to perform pixel-level extraction on the input satellite remote sensing image to output the roof mask, and calculate the original roof area based on the roof mask:
[0023] ;
[0024] in, For the first The original area of the roof This represents the actual area corresponding to a single pixel. For the roof mask in pixels The value at that location.
[0025] Furthermore, the multimodal fusion module includes:
[0026] Image branch features Used to receive the satellite remote sensing images and microclimate images, and extract deep image features of spatial texture, thermal environment and radiation environment from them;
[0027] Table branching features : Used to receive the tabular data and extract non-image tabular features of building attributes, spatial morphology and engineering parameters from it;
[0028] Fusion output layer: used to integrate the image branch features and the table branch features Mapping to the same feature space and fusing them to obtain a fused feature vector. :
[0029] ;
[0030] in, and For the mapping matrix, For bias terms, For activation function, This is the fused feature vector.
[0031] Furthermore, in the multimodal model construction and training steps, the loss function used for model training... for:
[0032] ;
[0033] in, For roof segmentation losses, To predict potential losses, and These are the weighting coefficients.
[0034] Furthermore, the multi-level potential assessment step further includes:
[0035] The installable area of the roof is calculated based on the original roof area, installation suitability coefficient, shading correction coefficient, and component layout utilization coefficient. The expression is as follows:
[0036] ;
[0037] in, For the first The area that can be installed on the roof. The original area of the roof. For the installation suitability factor, For occlusion correction factor, Component layout utilization factor;
[0038] The power generation potential of the technology is calculated based on the annual irradiance at the location of the roof, the installable area, the component efficiency, the environmental correction factor, and the overall system efficiency. The expression is as follows:
[0039] ;
[0040] in, For the first Annual technical power generation of a single rooftop This refers to the annual irradiance of the roof. For the rated efficiency of the component, This is the environmental correction factor. For overall system efficiency.
[0041] Furthermore, the multi-level potential assessment step also includes:
[0042] Based on the power generation potential of the technology, further calculate the carbon emission reduction potential:
[0043] ;
[0044] in, For the first Annual carbon emission reduction per rooftop Carbon emission factor per unit of electricity in the regional power grid;
[0045] Economic potential is calculated by considering the benefits and costs over the project's lifespan, and its expression is as follows:
[0046] ;
[0047] in, For the first The economic potential of a rooftop For initial investment costs, For the first Annual maintenance costs For the first Annual unit electricity price or alternative electricity price The annual decay rate, For the discount rate, The project's lifespan.
[0048] Furthermore, the method also includes:
[0049] The output results of the multimodal deep learning model are interpreted and verified. The interpretation includes analyzing the contribution of different data modalities to the prediction results through ablation experiments.
[0050] ;
[0051] in, For the first The contribution of each modality For complete model performance metrics, To remove the first Model performance metrics after each modality.
[0052] Furthermore, the verification includes:
[0053] For the roof identification results, the intersection-union ratio (IUU) is used to evaluate the segmentation accuracy:
[0054] ;
[0055] in, To correctly identify the number of pixels as the roof, The number of pixels misidentified as rooftops The number of roof pixels that were missed in identification;
[0056] For potential prediction results, the root mean square error is used to evaluate the prediction accuracy of continuous variables:
[0057] ;
[0058] in, For predicted values, For the true value, This represents the number of samples.
[0059] Secondly, the present invention provides a system for assessing urban photovoltaic potential based on multimodal deep learning, the system being used to execute the method, the system comprising:
[0060] The data acquisition and preprocessing unit is used to acquire and preprocess multimodal data from the study area to construct a unified multimodal input dataset.
[0061] The model building and training unit is used to build and train a multimodal deep learning model, which includes a roof recognition module for recognizing building roofs and a multimodal fusion module for fusing multimodal data to output photovoltaic potential correction parameters.
[0062] A multi-level potential assessment unit is used to calculate radiation potential, installation potential, technical power generation potential, carbon emission reduction potential and economic potential in sequence based on the roof mask output by the roof identification module and the photovoltaic potential correction parameters output by the multi-modal fusion module.
[0063] The result output unit is used to output a comprehensive potential zoning map, grading map, or ranking result that includes the multi-level potential assessment results.
[0064] Beneficial effects:
[0065] 1. High model performance: From single-modal image recognition to multi-modal correction, improving recognition accuracy and rationality.
[0066] Existing methods are usually based on a single remote sensing image, which can only identify the roof outline or estimate the roof area. They are difficult to consider factors such as building attributes, surrounding obstruction and microclimate environment at the same time, so they are prone to biases such as "identifiable but not necessarily installable" and "large area but low actual power generation capacity".
[0067] This invention, based on rooftop identification, simultaneously incorporates satellite imagery, microclimate imagery, building attribute data, and spatial morphology data. Through the joint expression of image and tabular features, it comprehensively corrects for rooftop installation suitability, shading relationships, and environmental efficiency. This expands the model output from a single geometric identification result to a comprehensive judgment oriented towards actual photovoltaic development conditions. In implementation cases, the error between the urban density identified by this method and the actual urban density is 2.9649%, and the error between the solar radiation result obtained by this method and the comparative simulation result is 9.51%, indicating that this method has a good accuracy foundation in both the two key aspects of rooftop identification and photovoltaic resource estimation.
[0068] 2. Fast computation speed: Improves the ability to conduct batch assessments at the city scale through multimodal integrated processing.
[0069] While existing single-modal deep learning methods can quickly identify rooftops, their outputs typically require subsequent steps such as building attribute screening, shading correction, and environmental correction to produce results usable for photovoltaic potential assessment, making the overall process quite fragmented. This invention integrates rooftop identification, multimodal condition correction, and potential calculation into a unified technical process, simultaneously obtaining key parameters such as installation suitability, shading correction, and environmental correction during the model output stage. This reduces subsequent step-by-step screening, repetitive overlay analysis, and fragmented computation steps.
[0070] Therefore, the advantages of this invention are not simply reflected in the reduced time for single model inference, but mainly in the improved batch processing capability and overall evaluation efficiency at the city scale while enhancing the completeness of the results. In the implementation case, the study area was divided into 1km×1km grids, forming a total of 31×31=961 computational units, verifying that this method can support the unified evaluation of large-scale spatial units.
[0071] 3. Strong engineering practicality: From single-dimensional area calculation to multi-dimensional comprehensive potential assessment for engineering applications, improving its feasibility.
[0072] Existing technologies often only provide results at the level of roof area or solar radiation, and the output is often a single resource value. They are difficult to answer key questions in engineering applications, such as which roofs are suitable for installation, how many can be installed, how much electricity can be generated, how much carbon emissions can be reduced, and whether it is economically feasible.
[0073] Building upon the model output, this invention further establishes a multi-level potential assessment system, including at least five levels: radiation potential, installation potential, technical power generation potential, carbon emission reduction potential, and economic potential. This allows the assessment results to directly serve urban energy planning, distributed photovoltaic layout, and project investment decisions. Case studies demonstrate that this method not only outputs overall potential but also differentiates development variations across different regional types. For example, the annual power generation results for different regional types reached 65.19 GWh, 63.16 GWh, and 73.03 GWh, respectively, indicating that the method's output is not merely a single resource identification value but rather an engineering result with spatial variability and decision-making reference value. Attached Figure Description
[0074] Figure 1 This is a flowchart illustrating a method for assessing urban photovoltaic potential based on multimodal deep learning, as proposed in an embodiment of the present invention.
[0075] Figure 2 This is a schematic diagram showing the photovoltaic utilization potential per square kilometer in a certain region. Detailed Implementation
[0076] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0077] Existing technologies for assessing the potential of urban building rooftop photovoltaics (PV) systems suffer from at least three shortcomings: First, they lack a multi-source information fusion mechanism, making it difficult to accurately characterize the conditions for rooftop PV development in complex urban environments. Second, they struggle to balance the speed and accuracy of urban-scale assessments; fast methods yield coarse results, while high-precision methods are often inefficient. Third, they lack a multi-level potential assessment system for engineering applications, limiting the direct application value of assessment results in urban energy planning, project selection, and investment decisions. Therefore, this invention provides a multi-modal deep learning-based method for assessing urban PV potential. This method addresses the task of assessing the development potential of urban building rooftop PV systems, constructing a technical framework consisting of basic multi-source data collection and processing, multi-modal model architecture and training, and multi-level potential assessment. The core idea is to first acquire and process image and tabular data related to urban building rooftops, then extract rooftop spatial features, environmental features, and attribute features through a multi-modal deep learning model, and finally combine the model output results to complete the hierarchical calculation of rooftop PV potential.
[0078] refer to Figure 1 and Figure 2 The overall technical process of this invention includes the following three sections:
[0079] First, basic multi-source data collection and processing. This involves acquiring high-resolution satellite remote sensing images, microclimate images, building attribute data, spatial morphology data, and engineering parameter data required for photovoltaic assessment of the study area. Image data is cropped, resampled, registered, and normalized; tabular data is cleaned, coded, and standardized to form a unified multimodal input dataset. This section provides a spatially consistent and formatted data foundation for subsequent model training and potential assessment.
[0080] Second, multimodal model architecture and training. A roof recognition model is built based on satellite imagery to extract the building roof boundary and original area. Then, a multimodal fusion model combining image branch and table branch is constructed. The image branch is used to extract deep features from satellite imagery, microclimate imagery, and building morphology imagery, while the table branch is used to extract building attributes and spatial morphology indicators. The two types of features are fused to output roof installation suitability, occlusion correction, environmental efficiency correction, and potential level or potential parameters.
[0081] Third, multi-level potential assessment. Based on the roof identification results and the output of the multimodal model, the roof's radiation potential, installation potential, technical power generation potential, carbon reduction potential, and economic potential are further calculated. Among them, radiation potential characterizes the solar energy resources available on the roof, installation potential characterizes the area where photovoltaic modules can be actually installed, technical power generation potential characterizes the theoretical power generation capacity under current module efficiency and system conditions, and carbon reduction potential and economic potential characterize its environmental and economic benefits, respectively.
[0082] (1) Basic multi-source data collection and processing.
[0083] This section serves as the data foundation for assessing the potential of rooftop photovoltaic (PV) systems on urban buildings. Input data includes three categories: image data, tabular data, and engineering parameter data. Image data is used to represent the spatial distribution characteristics of the rooftop and its surrounding environment; tabular data is used to represent building attributes and spatial statistical characteristics; and engineering parameter data is used to support the calculation of subsequent power generation, carbon emission reductions, and economic benefits.
[0084] Image data includes satellite remote sensing images and microclimate images. Satellite remote sensing images are used to identify building roof boundaries, locations, and areas; microclimate images are used to represent the spatial distribution of environmental factors such as surface temperature, solar radiation, wind speed, and humidity. Tabular data includes building age, building function, building height, roof type, roof slope, shading statistics, and socioeconomic data such as population density, age distribution, and income level. Engineering parameter data includes component efficiency, temperature coefficient, overall system efficiency, electricity price, and grid emission factors.
[0085] After data acquisition, data from different sources are processed uniformly. Image data undergoes coordinate transformation, resampling, cropping, and normalization sequentially; tabular data undergoes missing value cleanup, category coding, and numerical standardization sequentially; different modalities are aligned using spatial units or building objects, ensuring that the same building roof simultaneously corresponds to image information, attribute information, and environmental information. After the above processing, each sample can be represented as:
[0086] ;
[0087] in, For the first Satellite images of a sample, For the first A collection of microclimate images of individual samples. For the first The table feature vector of each sample.
[0088] For image data, normalization is used to eliminate differences in the numerical ranges of different channels. The expression is as follows:
[0089] ;
[0090] in, For the first The first sample The original values of each channel, and These are the minimum and maximum values for the corresponding channels, respectively.
[0091] For tabular data, numerical variables are standardized, and categorical variables are converted to discrete or one-hot encoding. After processing, the image modalities and tabular modalities are combined to form a multimodal sample set, which is then divided into training, validation, and test sets for subsequent model training and accuracy verification.
[0092] (2) Multimodal model architecture and training.
[0093] This section extracts roof identification results and potential correction parameters from multimodal input data. The model consists of two parts: a roof identification module and a multimodal fusion module. The former obtains the roof's geometric boundaries and original area, while the latter combines image and tabular features to output the intermediate parameters required for photovoltaic potential assessment.
[0094] The roof recognition module takes satellite imagery as input and employs a semantic segmentation network to extract building rooftops pixel-by-pixel, outputting a roof mask. Based on the number of roof pixels in the mask, the original roof area is calculated. Its area expression is:
[0095] ;
[0096] in, For the first The original area of the roof, This represents the actual area corresponding to a single pixel. For the roof mask in pixels The value at this location. This module outputs the roof boundary and the original geometric area, providing the basic quantities for subsequent condition adjustments.
[0097] The multimodal fusion module includes an image branch, a table branch, and a fusion output layer. The image branch receives satellite and microclimate images, extracting spatial texture, thermal environment, and radiation environment features; the table branch receives building attributes, spatial morphology, and engineering parameters, extracting non-image features; the fusion output layer maps the two types of features to the same feature space and performs joint expression to obtain the installation suitability coefficient, shading correction coefficient, environmental correction coefficient, and photovoltaic potential output. Its fusion expression is:
[0098] ;
[0099] in, For image branch features, For table branching features, and For the mapping matrix, For bias terms, For activation function, This is the fused feature vector.
[0100] During model training, multimodal samples are input into the network. After forward propagation, roof identification and potential prediction results are obtained. The loss is then calculated based on the labels, and the parameters are updated. If the output is a continuous value, mean squared error loss is used; if the output is a rank category, cross-entropy loss is used. The overall loss can be written as:
[0101] ;
[0102] in, For roof segmentation losses, To predict potential losses, and These are the weighting coefficients. The model parameters are updated iteratively through backpropagation until the validation set error meets the stopping condition.
[0103] After training, the model outputs the roof mask, original area, installation suitability coefficient, shading correction coefficient, environmental correction coefficient, and photovoltaic potential parameters for each roof.
[0104] (3) Multi-level potential assessment.
[0105] This section performs tiered calculations of the photovoltaic potential of urban building rooftops based on rooftop identification results and multimodal model outputs. The assessment tiers include radiation potential, installation potential, technical power generation potential, carbon emission reduction potential, and economic potential. Its purpose is to transform the rooftop area, installation suitability, shading correction, and environmental correction results output by the model into quantitative indicators that can be directly used for engineering decision-making.
[0106] First, calculate the installable area of the roof based on the original roof area, installation suitability coefficient, shading correction coefficient, and component layout utilization coefficient. The expression is:
[0107] ;
[0108] in, For the first The available installation area on the roof The original area of the roof. For the installation suitability factor, For occlusion correction factor, This is the component layout utilization factor. This formula is used to convert the geometric area into the area that can be arranged.
[0109] Secondly, the power generation potential of the technology is calculated based on the annual irradiance at the roof location, the installable area, the module efficiency, the environmental correction factor, and the overall system efficiency. The expression is:
[0110] ;
[0111] in, For the first Annual technical power generation of a single rooftop This refers to the annual irradiance of the roof. For the rated efficiency of the component, This is the environmental correction factor. This formula represents the overall system efficiency. It simultaneously reflects the relationship between three levels: radiation potential, installation potential, and technological potential: irradiance determines the resource base, the installable area determines the layout scale, and efficiency parameters determine the final power generation capacity.
[0112] Based on the technical power generation potential, the carbon emission reduction potential is further calculated. Its expression is:
[0113] ;
[0114] in, For the first Annual carbon emission reduction per rooftop This represents the carbon emission factor per unit of electricity generated by the regional power grid. This result is used to reflect the environmental benefits of photovoltaic power generation replacing conventional electricity.
[0115] Economic potential is calculated using the project's revenue and costs over its lifespan. Its expression is:
[0116] ;
[0117] in, For the first The economic potential of a rooftop For initial investment costs, For the first Annual maintenance costs For the first Annual unit electricity price or alternative electricity price The annual decay rate, For the discount rate, This represents the project's lifespan. This formula is used to evaluate the economic feasibility of a photovoltaic system throughout its entire lifespan.
[0118] For city-scale results, the corresponding indicators for all rooftops are summed to obtain the city's total radiation potential, total installation potential, total technical power generation potential, total carbon emission reduction potential, and total economic potential. After standardizing each indicator, a comprehensive potential zoning map, ranking map, or sorting result can be generated for determining the allocation of urban photovoltaic resources and the timing of development.
[0119] (4) Interpretation and verification of results.
[0120] This section interprets the model output and verifies the accuracy of roof identification and potential prediction results. Interpretation methods include ablation experiments, feature contribution analysis, and image response region analysis; verification methods include segmentation accuracy evaluation and potential prediction error evaluation. Its purpose is to illustrate the impact of different modalities and feature variables on the results and to verify whether the model output meets application requirements.
[0121] First, ablation experiments were used to analyze the contribution of different data modes to the prediction results. Specifically, satellite image modes, microclimate image modes, or tabular modes were removed sequentially, and the changes in model performance were compared. The contribution of a particular mode can be expressed as:
[0122] ;
[0123] in, For the first The contribution of each modality For complete model performance metrics, To remove the first Model performance metrics after each modality. If A larger value indicates that the mode has a significant impact on the photovoltaic potential prediction results.
[0124] Secondly, feature contribution analysis is used on the table features to identify the direction and degree of influence of building attributes, spatial morphology, and environmental variables on the model output; response heatmap analysis is used on the image branches to locate the areas in the image most sensitive to potential prediction. Through the above interpretation process, key factors affecting rooftop photovoltaic potential can be identified, such as roof area, building shading, surface temperature, and irradiance distribution, and output in the form of ordinal maps or heatmaps. For the roof identification results, the intersection-union ratio (IUU) is used to evaluate the segmentation accuracy. Its expression is:
[0125] ;
[0126] in, To correctly identify the number of pixels as the roof, The number of pixels misidentified as rooftops This represents the number of roof pixels that were missed. This metric is used to evaluate the consistency between the roof boundary extraction results and the true labels. For potential prediction results, the root mean square error (RMSE) is used to evaluate the prediction accuracy of continuous variables. Its expression is:
[0127] ;
[0128] in, For predicted values, For the true value, This represents the sample size. If the output is a potential level category, then precision, recall, and... The above indicators can be used to evaluate the results of the roof recognition module and the multimodal fusion module, respectively.
[0129] In a specific implementation scenario, the following study uses a particular region as a case study, as shown in Table 1:
[0130] Table 1. Case Data Sources and Definitions
[0131]
[0132] The calculation of case errors is shown in Tables 2 and 3 below:
[0133] Table 2 Algorithm Errors
[0134]
[0135] Table 3 Deep Learning Recognition Errors
[0136]
[0137] Case analysis results as follows Figure 2 As shown in the figure, the photovoltaic utilization potential per square kilometer in this region (×106 kWh / km2 / year) is as follows: (a: Power generation of industrial area 65.19 GWh; b: Power generation of high-density residential area 63.16 GWh; c: Power generation of industrial area 73.03 GWh).
[0138] Example embodiments have been disclosed herein, and while specific terminology has been used, it is for illustrative purposes only and should be construed as such, and is not intended to be limiting. In some instances, it will be apparent to those skilled in the art that features, characteristics, and / or elements described in conjunction with particular embodiments may be used alone, or in combination with features, characteristics, and / or elements described in conjunction with other embodiments, unless otherwise expressly indicated. Therefore, those skilled in the art will understand that various changes in form and detail may be made without departing from the scope of the invention as set forth in the appended claims.
Claims
1. A multi-modal deep learning based urban photovoltaic potential assessment method, characterized in that, include: S1. Acquire multimodal data of the study area. The multimodal data includes at least satellite remote sensing images, microclimate images, and tabular data. Preprocess the multimodal data to construct a unified multimodal input dataset. S2, Construct a multimodal deep learning model, the multimodal deep learning model including a roof recognition module for recognizing building roofs from the satellite remote sensing image, and a multimodal fusion module for fusing the multimodal data; The multimodal deep learning model is trained using the multimodal input dataset, and the trained model is used to output a roof mask and at least one photovoltaic potential correction parameter. S3. Based on the roof mask and the photovoltaic potential correction parameters, according to the preset hierarchical logic, the radiation potential representing solar energy resources, the installation potential representing the area where photovoltaic modules can be arranged, the technical power generation potential representing the theoretical power generation capacity, the carbon emission reduction potential representing environmental benefits, and the economic potential representing economic benefits are calculated in sequence to form a multi-level potential assessment result.
2. The method of claim 1, wherein, S1 specifically includes: The satellite remote sensing images and microclimate images are subjected to coordinate transformation, resampling, cropping and normalization to obtain image modal data; Missing value cleanup, category coding, and numerical standardization are performed on the table data to obtain table modal data, which includes building attribute data, spatial morphology data, and engineering parameter data; aligning the image modality data and the tabular modality data by spatial cells or building objects, so that each roof sample to be evaluated corresponds to a set of multi-modality data, the multi-modality data is represented as: ; wherein, is a satellite image for the th sample, is a set of microclimate images for the th sample, is a tabular feature vector for the th sample.
3. The method of claim 1, wherein, The roof identification module in S2 is a semantic segmentation network used to extract pixels from the input satellite remote sensing image to output the roof mask, and calculate the original roof area based on the roof mask. ; in, For the first The original area of the roof This represents the actual area corresponding to a single pixel. For the roof mask in pixels The value at that location.
4. The method according to claim 3, characterized in that, The multimodal fusion module includes: Image branch features Used to receive the satellite remote sensing images and microclimate images, and extract deep image features of spatial texture, thermal environment and radiation environment from them; Table branching features : Used to receive the tabular data and extract non-image tabular features of building attributes, spatial morphology and engineering parameters from it; Fusion output layer: used to integrate the image branch features and the table branch features Mapping to the same feature space and fusing them to obtain a fused feature vector. : ; in, and For the mapping matrix, For bias terms, For activation function, This is the fused feature vector.
5. The method according to claim 4, characterized in that, In the multimodal deep learning model construction and training steps, the loss function used for model training is... for: ; in, For roof segmentation losses, To predict potential losses, and These are the weighting coefficients.
6. The method according to claim 5, characterized in that, The multi-level potential assessment steps further include: The installable area of the roof is calculated based on the original roof area, installation suitability coefficient, shading correction coefficient, and component layout utilization coefficient. The expression is as follows: ; in, For the first The area that can be installed on the roof. The original area of the roof. For the installation suitability factor, For occlusion correction factor, Component layout utilization factor; The power generation potential of the technology is calculated based on the annual irradiance at the location of the roof, the installable area, the component efficiency, the environmental correction factor, and the overall system efficiency. The expression is as follows: ; in, For the first Annual technical power generation of a single rooftop This refers to the annual irradiance of the roof. For the rated efficiency of the component, This is the environmental correction factor. For overall system efficiency.
7. The method according to claim 6, characterized in that, The multi-level potential assessment steps also include: Based on the power generation potential of the technology, further calculate the carbon emission reduction potential: ; in, For the first Annual carbon emission reduction per rooftop Carbon emission factor per unit of electricity in the regional power grid; Economic potential is calculated by considering the benefits and costs over the project's lifespan, and its expression is as follows: ; in, For the first The economic potential of a rooftop For initial investment costs, For the first Annual maintenance costs For the first Annual unit electricity price or alternative electricity price The annual decay rate, For the discount rate, The project's lifespan.
8. The method according to claim 1, characterized in that, The method further includes: The output results of the multimodal deep learning model are interpreted and verified. The interpretation includes analyzing the contribution of different data modalities to the prediction results through ablation experiments. ; in, For the first The contribution of each modality For complete model performance metrics, To remove the first Model performance metrics after each modality.
9. The method according to claim 8, characterized in that, The verification includes: For the roof identification results, the intersection-union ratio (IUU) is used to evaluate the segmentation accuracy: ; in, To correctly identify the number of pixels as the roof, The number of pixels misidentified as rooftops The number of roof pixels that were missed in identification; For potential prediction results, the root mean square error is used to evaluate the prediction accuracy of continuous variables: ; in, For predicted values, For the true value, This represents the number of samples.
10. A system for assessing urban photovoltaic potential based on multimodal deep learning, characterized in that, The system is used to perform the method according to any one of claims 1-9, the system comprising: The data acquisition and preprocessing unit is used to acquire and preprocess multimodal data from the study area to construct a unified multimodal input dataset. The model building and training unit is used to build and train a multimodal deep learning model, which includes a roof recognition module for recognizing building roofs and a multimodal fusion module for fusing multimodal data to output photovoltaic potential correction parameters. A multi-level potential assessment unit is used to calculate radiation potential, installation potential, technical power generation potential, carbon emission reduction potential and economic potential in sequence based on the roof mask output by the roof identification module and the photovoltaic potential correction parameters output by the multi-modal fusion module. The result output unit is used to output a comprehensive potential zoning map, grading map, or ranking result that includes the multi-level potential assessment results.