Soil nutrient prediction method for photovoltaic power station based on multi-source remote sensing data fusion

By integrating multi-source remote sensing data and using machine learning models, the problem of low accuracy in soil nutrient prediction for photovoltaic power plants has been solved, achieving efficient and accurate soil nutrient analysis and prediction, which is applicable to ecological restoration and vegetation planting planning for photovoltaic power plants.

CN121721253BActive Publication Date: 2026-05-01NORTHWEST INST OF ECO ENVIRONMENT & RESOURCES CAS
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NORTHWEST INST OF ECO ENVIRONMENT & RESOURCES CAS
Filing Date
2026-02-27
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies are difficult to apply effectively to soil nutrient prediction for photovoltaic power plants, especially in environments with significant spatial differentiation caused by the deployment of photovoltaic modules. The lack of comprehensive utilization of multi-source remote sensing data leads to low prediction accuracy and limited applicability.

Method used

A multi-source remote sensing data fusion method was adopted to acquire multispectral, thermal infrared and three-dimensional remote sensing data by UAV. After data preprocessing, the microenvironment of the photovoltaic power station was identified, the area under the panel, the front edge of the panel and the area between the panels were divided, soil sampling points were set up, and soil nutrient prediction was generated by training a machine learning model.

Benefits of technology

It improves the relevance and reliability of soil nutrient analysis, significantly enhances prediction accuracy, reduces detection costs, and enables rapid analysis of the spatial distribution of soil nutrients within photovoltaic power plants, providing a scientific basis for ecological restoration and vegetation planting planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121721253B_ABST
    Figure CN121721253B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of ecological environment monitoring of photovoltaic power station, and discloses a soil nutrient prediction method for photovoltaic power station based on multi-source remote sensing data fusion, which comprises the following steps: obtaining and preprocessing multispectral, thermal infrared and three-dimensional remote sensing data of a target photovoltaic power station to generate orthophoto, surface temperature distribution map and digital surface model; extracting road area to construct an effective analysis area mask, dividing three types of micro-environment (under the photovoltaic panel, in front of the photovoltaic panel and between the photovoltaic panels) by extracting the photovoltaic panel projection and the front edge of the photovoltaic panel; arranging soil sampling points in the micro-environment, determining soil nutrients as true values, and synchronously generating multi-source remote sensing features; training a machine learning model with remote sensing features as input and soil nutrient true values as output to generate a trained machine learning model; finally, inputting multi-source remote sensing features of a photovoltaic power station to be tested can predict the soil nutrient results; the present application can efficiently predict the soil nutrients and their spatial distribution in different micro-environmental areas of the photovoltaic power station.
Need to check novelty before this filing date? Find Prior Art

Description

Soil Nutrient Prediction Method for Photovoltaic Power Plants Based on Multi-Source Remote Sensing Data Fusion Technical Field

[0001] This invention relates to the field of ecological environment monitoring technology for photovoltaic power plants, specifically to a method for predicting soil nutrients in photovoltaic power plants based on multi-source remote sensing data fusion. Background Technology

[0002] The tilted installation structure and array arrangement of photovoltaic modules alter the redistribution of sunlight, moisture, and heat on the Earth's surface, leading to significant spatial differentiation in sunlight, temperature, and moisture conditions within the power plant, thus creating a photovoltaic microenvironment with strong spatial heterogeneity. Specifically, this manifests as follows:

[0003] (1) The area under the photovoltaic panel is in a state of long-term shade and rain protection, and the input of light and precipitation is significantly reduced, forming a shading environment with low light and low evaporation.

[0004] (2) The photovoltaic panels are completely exposed to solar radiation, forming a high-light and high-evaporation exposure environment.

[0005] (3) The area in front of the photovoltaic panel becomes a water catchment area where water infiltrates due to the convergence of runoff.

[0006] The spatially heterogeneous photovoltaic microenvironment further drives differences in soil and vegetation conditions within the power station, ultimately leading to significant spatial differentiation of soil nutrients within the power station.

[0007] Traditional soil nutrient testing relies on intensive field sampling and laboratory chemical analysis, which is costly, inefficient, and struggles to capture the complex spatial variations within photovoltaic power plants. Unmanned aerial vehicle (UAV) remote sensing technology, with its high spatiotemporal resolution and flexibility, offers the possibility of rapidly acquiring surface information over large areas. The multispectral sensors onboard UAVs can acquire the reflectance characteristics of vegetation canopies and exposed soil, and can indirectly characterize soil moisture and nutrient status by constructing spectral indices.

[0008] However, current solutions for applying UAV remote sensing technology to soil nutrient prediction are mainly aimed at natural farmland or areas with relatively uniform underlying surface conditions. Existing technologies are difficult to apply effectively to photovoltaic power station environments where soil nutrients exhibit significant spatial differentiation due to the deployment of photovoltaic modules. Furthermore, existing methods still have shortcomings in the comprehensive utilization of multi-source remote sensing data. There is a lack of methods capable of characterizing the microenvironmental features of photovoltaic power stations, making it difficult to establish a stable and reliable mapping model between multi-spectral information, terrain parameters, and other multi-source remote sensing features acquired by UAVs and soil nutrient indicators. Consequently, the accuracy and applicability of soil nutrient prediction results in photovoltaic power station scenarios are insufficient. Summary of the Invention

[0009] To address the aforementioned shortcomings in existing technologies, this invention provides a method for predicting soil nutrients in photovoltaic power plants based on multi-source remote sensing data fusion, thereby solving the problem of low prediction accuracy of soil nutrients in photovoltaic power plants.

[0010] To achieve the above-mentioned objectives, the technical solution adopted by this invention is as follows:

[0011] A method for predicting soil nutrients in photovoltaic power plants based on multi-source remote sensing data fusion includes the following steps:

[0012] Acquire multispectral, thermal infrared, and three-dimensional remote sensing data of the target photovoltaic power station, perform data preprocessing, and generate orthophotos, surface temperature distribution maps, and digital surface models.

[0013] Based on orthophotos and digital surface models, road areas are identified and extracted, and a mask for effective analysis areas excluding roads is generated.

[0014] Within the effective analysis area mask, the photovoltaic panel projection area and its front edge are extracted. By calculating the symbolic horizontal distance, three types of microenvironments are divided: the area under the panel, the front edge area of ​​the panel, and the area between the panels.

[0015] Soil sampling points were set up in three types of microenvironments to collect soil samples and determine their nutrient content. These samples were used as true soil nutrient data, and multi-source remote sensing features were generated.

[0016] Using multi-source remote sensing features as input and soil nutrient real data as output, the machine learning model is trained to generate a trained machine learning model.

[0017] The system acquires multi-source remote sensing features of the photovoltaic power station under test, inputs them into a trained machine learning model, and generates soil nutrient results for the photovoltaic power station under test.

[0018] The present invention has the following beneficial effects:

[0019] The proposed method for predicting soil nutrients in photovoltaic power plants based on multi-source remote sensing data fusion, through multi-source remote sensing data fusion and preprocessing, combined with a road area exclusion mechanism, accurately locks the effective analysis range, avoids interference from irrelevant areas, and improves the pertinence and reliability of soil nutrient analysis. Furthermore, by classifying and deploying soil sampling points according to microenvironmental categories, the true soil nutrient data more closely matches the actual scenario of photovoltaic power plants, providing high-quality samples for model training, significantly improving the prediction accuracy of machine learning models, and efficiently predicting the spatial distribution of soil nutrients in different microenvironmental areas within photovoltaic power plants. Simultaneously, it eliminates the need for large-scale on-site sampling, achieving rapid analysis through the combination of multi-source remote sensing features and models, significantly reducing detection costs and improving efficiency. It can also quickly output the soil nutrient results of the photovoltaic power plant under test, providing a scientific basis for ecological restoration and vegetation planting planning around photovoltaic power plants, possessing strong practicality and promotional value. Attached Figure Description

[0020] Figure 1 is a flowchart illustrating the method for predicting soil nutrients in photovoltaic power plants based on multi-source remote sensing data fusion proposed in this invention. Detailed Implementation

[0021] The specific embodiments of the present invention are described below to enable those skilled in the art to understand the present invention. However, it should be understood that the present invention is not limited to the scope of the specific embodiments. For those skilled in the art, various changes are obvious as long as they are within the spirit and scope of the present invention as defined and determined by the appended claims. All inventions utilizing the concept of the present invention are protected.

[0022] As shown in Figure 1, the method for predicting soil nutrients in photovoltaic power plants based on multi-source remote sensing data fusion includes the following steps:

[0023] Acquire multispectral, thermal infrared, and three-dimensional remote sensing data of the target photovoltaic power station, perform data preprocessing, and generate orthophotos, surface temperature distribution maps, and digital surface models.

[0024] In this embodiment, this step involves using a drone to synchronously acquire multi-source remote sensing data of the target photovoltaic power station. The operation process is as follows:

[0025] A high-precision positioning UAV equipped with a multispectral camera, a thermal infrared camera, and a 3D information acquisition device executes a flight mission according to preset flight parameters to acquire multispectral, thermal infrared, and 3D remote sensing data of the target photovoltaic power station, thus obtaining multi-source remote sensing data.

[0026] In this embodiment, firstly, the flight and mission payload are configured; a high-precision positioning UAV equipped with a real-time dynamic differential (RTK) or post-processing dynamic differential (PPK) module is selected to achieve centimeter-level positioning accuracy and ensure accurate spatial registration of multi-source remote sensing data; and this high-precision positioning UAV also needs to be equipped with the following three types of sensors simultaneously:

[0027] (1) Multispectral camera; its spectral bands must cover blue light with a center wavelength of about 450 nm, green light with a center wavelength of about 560 nm, red light with a center wavelength of about 650 nm, red edge with a center wavelength of 710 nm - 740 nm, and near-infrared light with a center wavelength of about 840 nm; the multispectral camera is used to obtain the surface spectral reflectance information of photovoltaic panels in photovoltaic power stations, so as to calculate vegetation index and soil index in subsequent steps.

[0028] (2) Thermal infrared camera; its thermal sensitivity (NETD) should be less than 0.1K, and its spatial resolution is comparable to that of multispectral data; this thermal infrared camera is used to directly obtain the surface brightness and temperature information of photovoltaic panels in photovoltaic power plants, which is the key to quantifying the spatial temperature difference caused by the shading effect of photovoltaic panels.

[0029] (3) Three-dimensional information acquisition equipment; it is used to accurately acquire the three-dimensional structural information of photovoltaic panels in photovoltaic power stations, and can be measured by optical photogrammetry or lidar, specifically:

[0030] Optical photogrammetry; equipped with a high-resolution visible light camera with at least 20 million pixels; this method is cost-effective and suitable for scenes with low vegetation cover;

[0031] LiDAR measurement; equipped with a LiDAR scanner, the point cloud density should be no less than 50 points per square meter; this method can actively acquire high-precision three-dimensional point clouds, is less affected by light, and can effectively penetrate low vegetation, more accurately depicting the surface morphology, especially with significant advantages in vegetation-covered areas.

[0032] Secondly, the flight mission is executed according to the preset flight parameters to collect multi-source remote sensing data; the preset flight parameters are as follows:

[0033] (1) Ground resolution; the final ground sampling distance for all sensor data should be less than 5 cm; this resolution is a necessary condition for accurate identification and quantification of the board front area with a width of about 20 cm;

[0034] (2) Flight mode; To overcome the obstruction of vertical observation by photovoltaic panels, this invention adopts a hybrid flight path mode that combines vertical and tilted flight paths; The vertical flight path is that each camera sensor shoots vertically downwards at the ground to obtain orthophotos and three-dimensional models of the photovoltaic power station; The tilted flight path is that flight paths are set up on the north and south sides of the photovoltaic array to control each camera sensor to shoot at the photovoltaic panel array at a suitable angle; Among them, the north tilted flight path is crucial; Since the photovoltaic panels are laid tilted to the south, the space under the panels is mainly open to the north; By taking tilted photography from the north at a low angle, the sensor line of sight can directly penetrate the space under the panels, effectively and directly obtaining multispectral data (i.e., surface spectral reflectance information) and thermal infrared data (i.e., surface brightness and temperature information) of the area under the photovoltaic panels, thereby significantly reducing the blind spot of observation; The south flight path can be used as a supplement to obtain information on the front of the photovoltaic panels and make the three-dimensional model more complete.

[0035] (3) Flight time; Multispectral data acquisition should be carried out in clear, cloudless weather conditions with stable light and local time between 10:00 and 14:00; Thermal infrared data acquisition should be carried out in clear, windless conditions and local time between 13:00 and 15:00 to maximize the surface temperature difference between the area under the plate and the area between the plates.

[0036] (4) Route parameters; a grid-like route is adopted, with a heading overlap rate of not less than 80% and a lateral overlap rate of not less than 70%, to ensure the accuracy and completeness of the 3D modeling;

[0037] (5) Data synchronization: During a flight operation, all the above sensors are triggered synchronously, and the time system of all devices must be strictly synchronized with the UAV GNSS clock to ensure the acquisition of multi-source remote sensing data with a unified spatiotemporal reference.

[0038] Radiometric calibration, atmospheric correction, geometric correction, and orthorectification are performed on multispectral data to generate orthorectified images of the Earth's surface spectral reflectance.

[0039] In this embodiment, this step involves preprocessing the multispectral data, specifically as follows:

[0040] First, radiometric calibration is performed; using the calibration coefficients built into the sensor, the original digital quantization value ( This is converted into the radiance value of the upper atmosphere. This process can be expressed as: , Radiance, , All are scaling factors;

[0041] Secondly, atmospheric correction is performed; a method based on the radiative transfer model is used to eliminate the effects of atmospheric scattering and absorption, and the radiance of the top atmospheric layer is converted into the true reflectance of the Earth's surface.

[0042] Then, geometric correction and orthorectification are performed; based on the UAV POS system data and the generated high-precision digital surface model, the geometric distortion caused by camera tilt and terrain undulation is corrected by collinearity equations, generating an orthorectified image of the surface reflectance with accurate geographic coordinates, and finally realizing the data preprocessing of multispectral data.

[0043] Radiometric calibration, atmospheric correction, and geometric correction are performed on thermal infrared data to generate a surface temperature distribution map based on surface brightness and temperature information.

[0044] In this embodiment, this step involves preprocessing the thermal infrared data, specifically as follows:

[0045] First, radiometric calibration is performed; the raw digital values ​​are converted into radiance values ​​at the sensor.

[0046] Secondly, atmospheric correction and temperature inversion are performed. Based on the theory of atmospheric radiative transfer, considering the influence of atmospheric transmittance and path radiation, the sensor radiance value is converted into the surface radiance. Subsequently, according to the inverse function of Planck's blackbody radiation law, the surface radiance is converted into the true surface temperature (unit: degrees Celsius).

[0047] Then, geometric correction is performed; in conjunction with multispectral data, it is registered to the same geographic coordinate system to generate a surface temperature distribution map;

[0048] Determine whether the 3D remote sensing data is optical photogrammetry. If so, generate a digital surface model through feature point matching and interpolation. Otherwise, if it is lidar measurement, perform filtering, classification, and interpolation to generate a digital surface model.

[0049] In this embodiment, this step involves preprocessing the three-dimensional remote sensing data, specifically as follows:

[0050] If optical photogrammetry is used, based on multi-view images (including vertical and oblique images), the Structure of Motion (SfM) and Multi-View Stereo Matching (MVS) methods are employed. The mathematical foundation is the collinearity equation in photogrammetry. The exterior orientation elements of each image are automatically calculated through feature point matching, generating a densely matched point cloud. Finally, a high-precision digital surface model (DSM) is generated through interpolation. This equation expresses the strict geometric relationship between the object point, the projection center, and the image point, and its expression is:

[0051]

[0052]

[0053] in, For image point coordinates, For the camera's orientation elements, Let these be the coordinates of the camera's principal point. For camera focal length, Let these be the coordinates of the object point. The coordinates of the projection center are, , , All represent elements of a rotation matrix. The value can be 1 to 3;

[0054] The exterior orientation elements of each image are automatically calculated through feature point matching (such as the scale-invariant feature transform method SIFT and the accelerated robust feature method SURF). The rotation matrix), that is, this step is the core automated process of transforming a bunch of two-dimensional photos whose exact location and angle are unknown into a set of observation data with precise three-dimensional spatial coordinates and attitude, which is the geometric basis for all subsequent three-dimensional reconstruction (generating point clouds, DSM).

[0055] If LiDAR is used, filtering and classification are performed first, using a method based on point cloud geometric features (such as progressive triangulation densification filtering). This method identifies ground points by iteratively constructing a triangulation and setting a distance threshold. Its core discrimination criterion is:

[0056]

[0057] in, Let be the vertical distance from the point to be determined to the current triangulation network. To construct the elevation difference of the point set in the triangulation network, The scale factor is used to classify the original point cloud into ground points and non-ground points. Then, spatial interpolation methods such as irregular triangular meshes or inverse distance weighting are used to generate a digital surface model (DSM) from the classified point cloud. Specifically, interpolation using only ground points can further generate a digital terrain model (DTM) representing bare terrain.

[0058] Based on orthophotos and digital surface models, road areas are identified and extracted, and a mask for effective analysis areas excluding roads is generated.

[0059] In this embodiment, to ensure that the soil nutrient prediction model proposed in subsequent steps is only applicable to natural soil environments and to avoid non-soil target features, mainly to avoid interference from artificial roads in model training and prediction, road areas need to be identified and excluded before analysis; the operation process is as follows:

[0060] First, the road area is vectorized to obtain a precise vector polygon layer of the roads inside the target photovoltaic power station, specifically:

[0061] Based on the automatic extraction of remote sensing images; based on the above steps, orthophotos (DOM) and digital surface models (DSM) are generated. The road areas are automatically identified and vectorized using image segmentation and classification methods to obtain the road vector layer of the photovoltaic power station. The principle is that roads usually present regular linear or strip features on orthophotos, have uniform spectral reflectance characteristics that are significantly different from natural soil, and appear as continuous and flat surfaces on digital surface models.

[0062] Among them, the acquisition of the road vector map layer of photovoltaic power station can also be achieved through semi-automatic extraction of auxiliary data. Specifically, the design drawings and as-built drawings of photovoltaic power station are imported, which clearly indicate the road layout. After the above data is geometrically registered in geographic information system software, the accurate road vector polygon layer, i.e., the road vector map layer of photovoltaic power station, is generated through screen digitization or data format conversion.

[0063] Secondly, effective analysis region mask generation is performed, specifically as follows:

[0064] The road vector polygons obtained from the above steps are defined as non-analysis areas. Through spatial analysis, a complementary mask for effective analysis areas excluding roads is generated. This operation can be completed using ArcGIS software. This mask clarifies the geographical scope applicable to all subsequent analysis operations (including microenvironment quantification, soil sampling, and model prediction), namely, all natural soil areas within the power station excluding roads.

[0065] Within the effective analysis area mask, the photovoltaic panel projection area and its front edge are extracted. By calculating the symbolic horizontal distance, three types of microenvironments are divided: the area under the panel, the area at the front edge of the panel, and the area between the panels.

[0066] In this embodiment, to perform microenvironment quantization within the effective analysis region mask definition, the operation is as follows:

[0067] First, the projection area of ​​the photovoltaic panel and the front edge of the panel are accurately extracted, specifically as follows:

[0068] Within the effective analysis area mask, a planar segmentation method is used to identify the 3D point cloud of the photovoltaic panel from the digital surface model, project it vertically onto the horizontal plane, and generate several photovoltaic panel projection areas by calculating the 2D convex hull. The average elevation of each side of each photovoltaic panel projection area is compared, and the edge with the lowest elevation is taken as the front edge of the panel.

[0069] In this embodiment, a digital surface model is input, and a plane segmentation method is used to automatically identify photovoltaic panels from the digital surface model. This plane segmentation method, through iterative random sampling and model verification, can robustly fit the best plane model from complex scenes, and is particularly suitable for identifying large-area, flat, inclined planes presented by tilted photovoltaic panels. By repeating this process and clustering, a precise three-dimensional point cloud set for each photovoltaic panel can be obtained. Then, the point cloud of each photovoltaic panel is vertically projected onto a horizontal plane, and its two-dimensional convex hull is calculated to obtain the projection area of ​​the photovoltaic panel, which is a precise two-dimensional vector polygon. For each photovoltaic panel projection area, its south-facing edge is identified as the front edge of the panel. This identification is automatically completed by comparing the average elevation of each side of the projection area in three-dimensional space (the edge with the lowest elevation is the front edge of the panel).

[0070] Secondly, the determination of the dominant photovoltaic panel and the calculation of the symbolic horizontal distance are performed, specifically as follows:

[0071] For each pixel within the mask of the effective analysis region (Its horizontal projection coordinates are) ), perform the following sub-steps:

[0072] (1) Determine the dominant photovoltaic panel and adopt a layered judgment logic:

[0073] 1) Determining internal points: If a point... If a photovoltaic panel is located within any photovoltaic panel projection area, then the photovoltaic panel corresponding to that projection area is directly used as its dominant photovoltaic panel, specifically:

[0074] a) The pixel to be judged Projecting the image onto the two-dimensional plane containing the photovoltaic panel's projection area yields the projection point. ;

[0075] b) Iterate through the projected polygons of all photovoltaic panels in the two-dimensional plane;

[0076] c) For the currently traversed photovoltaic panel projection polygon, determine the projection points. Whether it is located inside the projected polygon can be determined using any of the following methods:

[0077] When the projected polygon is a convex polygon, construct the edge vectors and corresponding point vectors in sequence according to the vertex order, and calculate the two-dimensional cross product of the point vectors and each edge vector; if all cross product results have the same sign, then the projected point is determined. Located inside or on the boundary of the projected polygon;

[0078] When the projected polygon is a rectangle parallel to the coordinate axes, obtain the minimum and maximum coordinate values ​​of the rectangle in the X and Y directions to determine the projection point. If the coordinates of the points both fall within the corresponding coordinate interval, then determine the projection point. Located inside or on the boundary of the rectangle; otherwise, point Located outside the photovoltaic panel projection area, perform external point determination;

[0079] d) If the projection point is determined in step c) If a pixel is located inside or on the boundary of the current photovoltaic panel's projected polygon, then the photovoltaic panel is defined as a pixel. The dominant photovoltaic panel was selected, and the determination process was concluded.

[0080] 2) External point determination: If point If the point is not located within any photovoltaic panel projection area, then the point whose front edge is located is selected. All photovoltaic panels on the north side are considered as candidate photovoltaic panels, and calculation points are calculated. The candidate photovoltaic panel with the shortest vertical distance to the front edge of all candidate photovoltaic panels is selected as the dominant photovoltaic panel.

[0081] (2) Calculate the symbolic horizontal distance D

[0082] Calculation points The shortest Euclidean distance to the boundary of the dominant photovoltaic panel projection area is used as the absolute value of the symbolic horizontal distance (|D|) to determine the point. Whether it is located inside the projection area of ​​the dominant photovoltaic panel, if so, the symbolic horizontal distance is negative (D=-|D|), otherwise, the symbolic horizontal distance is positive (D=+|D|).

[0083] Finally, based on a preset distance threshold, the microenvironment types are finely classified, specifically as follows:

[0084] Based on the calculated symbolic horizontal distance D, all pixels are divided into three microenvironments, according to the following rules:

[0085] Under-panel region: defined as all pixels that satisfy D<0, assigned the classification label 1 (under-panel).

[0086] Frontal area of ​​the photovoltaic panel: defined as all pixels satisfying 0≤D≤δ, assigned classification label 2 (frontal area of ​​the panel). Here, δ is a preset distance threshold, which can be defined according to the boundary of vegetation growth along the frontal area of ​​the panel. In field surveys of vegetation and soil at photovoltaic power stations, it was found that within approximately 0.2 m outside the frontal area of ​​the photovoltaic panel, vegetation cover, plant height, and growth vigor show clear demarcation characteristics compared to the shaded area under the panel and the exposed area between panels. Furthermore, when the distance threshold δ is 0.2 m, approximately four consecutive pixels (0.2 m / 0.05 m = 4) can be covered along the frontal direction of the photovoltaic panel in the corresponding image, effectively reducing the impact of single-pixel noise, geometric registration error, and local surface inhomogeneity on the classification results.

[0087] Inter-plate region: defined as all pixels that satisfy D>δ, and assigned the classification label 3 (inter-plate).

[0088] Soil sampling points were set up in three types of microenvironments to collect soil samples and determine their nutrient content. These samples were used as true soil nutrient data, and multi-source remote sensing features were generated.

[0089] In this embodiment, the process of constructing a multi-source remote sensing feature dataset is as follows:

[0090] First, we need to obtain true soil nutrient data, specifically:

[0091] A stratified random sampling method was used to set up soil sampling points in three types of microenvironments and record the coordinates of each soil sampling point.

[0092] Soil samples were extracted from each soil sampling point to obtain the soil organic matter content, nitrogen content, phosphorus content, and potassium content of each soil sample, which were used as true soil nutrient data.

[0093] In this embodiment, within the photovoltaic power station, soil sampling points are deployed using a stratified random sampling method based on the pre-defined microenvironment types (under-panel area, front-panel area, and inter-panel area) to ensure sufficient and representative sample distribution in each microenvironment area. Simultaneously, the coordinates of each soil sampling point are recorded, with a positioning error of less than 0.05 meters. Soil samples from the surface layer of each sampling point are collected, processed according to standard procedures, and then sent to a qualified laboratory for chemical analysis. Following national standard methods, the soil organic matter content, nitrogen content, phosphorus content, and potassium content of each sample are accurately determined. These results will serve as irreplaceable ground truth values ​​for subsequent model training and validation.

[0094] Secondly, multi-source remote sensing feature extraction is performed, specifically as follows:

[0095] For each soil sampling point, combined with the coordinates of each soil sampling point, the original surface reflectance values ​​of blue light, green light, red light, red edge, and near-infrared bands at the coordinates of each soil sampling point are extracted from the corresponding orthophoto. Derived spectral indices sensitive to soil and vegetation attributes are calculated, including soil organic carbon index, normalized difference vegetation index, normalized difference vegetation index based on green light band, enhanced vegetation index, soil-modified vegetation index, and red edge chlorophyll index, which are used as spectral features.

[0096] In this embodiment, for each soil sampling point, based on its precise coordinates, a multi-dimensional feature vector is systematically extracted from its corresponding remote sensing products and quantification parameters. The extracted features cover spectral features, thermal infrared features, and microenvironmental geometric features. Among them, the spectral features are the original surface reflectance values ​​of blue, green, red, red-edge, and near-infrared bands at the coordinates of each soil sampling point extracted from orthophotos, and a series of derived spectral indices sensitive to soil and vegetation properties calculated based on these surface reflectance values. The calculation formulas for these indices are as follows:

[0097] Soil Organic Carbon Index :

[0098]

[0099] in, , , These represent the reflectance in the blue light band, red light band, and green light band, respectively. The reflectance needs to be scaled during calculation (for example, for Sentinel-2 data, the reflectance value is multiplied by 10000, i.e., the range of 0-10000 is used).

[0100] Normalized Difference Vegetation Index (NDVI) ):

[0101]

[0102] in, It represents the reflectance in the near-infrared band; and this normalized differential vegetation index is used to assess vegetation cover and biomass.

[0103] The Green Normalized Difference Vegetation Index (GFVI) is based on the green light band. :

[0104] ;

[0105] Enhance vegetation index

[0106] ;

[0107] Soil-modified vegetation index (SDI) ):

[0108]

[0109] in, This represents the soil conditioning coefficient, a constant introduced to reduce the influence of soil background. It is usually set to 0.5-1.0 and has a significant effect in areas with low vegetation cover.

[0110] Red-edged chlorophyll index ( ):

[0111]

[0112] in, It represents the reflectance of the red-edge band; and this red-edge chlorophyll index is very sensitive to the chlorophyll content of leaves, and can be used as an important indicator for inverting the nitrogen status of vegetation.

[0113] For each soil sampling point, the surface temperature value at each soil sampling point's coordinates is extracted from the corresponding surface temperature distribution map and used as a thermal infrared feature.

[0114] In this embodiment, the thermal infrared features are mainly the surface temperature values ​​(LST, °C) at the coordinates of each soil sampling point extracted from the surface temperature distribution map.

[0115] Meanwhile, the symbolic horizontal distance value and the classification labels corresponding to the three types of microenvironments are used as the geometric features of the microenvironment.

[0116] Ultimately, spectral features, thermal infrared features, and microenvironment geometric features were used as multi-source remote sensing features.

[0117] In this embodiment, spectral features, thermal infrared features, and microenvironment geometric features are used as multi-source remote sensing features to construct a dataset. Specifically, in a GIS or professional data processing environment, the unique coordinates of each soil sampling point are used to accurately connect and match all the above-mentioned multi-source remote sensing features with the ground truth data of soil nutrients, ensuring the uniqueness and accuracy of each data record. The final constructed dataset is represented as a structured two-dimensional table, where each row represents an independent soil sample and each column corresponds to a feature variable or target variable. The table includes at least the sample identifier (ID), geographic coordinates, all spectral feature values, surface temperature values, symbolic horizontal distance values, classification labels corresponding to the three types of microenvironments, and laboratory measured values ​​of four soil nutrients. To ensure the objectivity of model evaluation, a random sampling method is used to divide the complete dataset into a training subset and an independent validation subset. The training subset typically accounts for 70% to 80% of the total sample size and is used for model training and parameter tuning. The independent validation subset accounts for 20% to 30% and is specifically used to objectively evaluate the generalization performance and prediction accuracy of the trained model. Therefore, the dataset constructed based on multi-source remote sensing features successfully quantified the correlation between the apparent information obtainable by remote sensing and the intrinsic properties of the soil, laying a solid foundation for the next stage of data-driven model construction.

[0118] Using multi-source remote sensing features as input and soil nutrient data as output, a machine learning model is trained to generate a well-trained machine learning model.

[0119] Specifically, the machine learning model is either a random forest model or a gradient boosting decision tree model.

[0120] In this embodiment, an ensemble learning method that can effectively handle high-dimensional features, capture complex nonlinear relationships, and is insensitive to multicollinearity is selected as the core modeling framework. Therefore, this invention selects a random forest model or a gradient boosting decision tree model, trains it by inputting multi-source remote sensing features, and generates a trained machine learning model. This trained model is then used as a soil nutrient prediction model to predict soil nutrients.

[0121] Random forest models, which construct and integrate multiple decision trees, have the advantages of high training efficiency, low overfitting, and the ability to provide feature importance ranking.

[0122] The core of the gradient boosting decision tree model lies in iteratively training a series of decision trees. Each tree learns to correct the residuals of the previous tree, and finally, a strong predictor is obtained by weighted summation. This model typically achieves higher prediction accuracy. Taking the XGBoost method as an example, its prediction model can be expressed as:

[0123]

[0124] in, Indicates the first training subset The predicted value for each sample, This represents a gradient boosting decision tree model. Indicates the first Multi-source remote sensing feature vectors of each sample This represents the total number of trees. Indicates the first A decision tree.

[0125] Its objective function typically includes a loss function. With regularization term ,Right now:

[0126]

[0127]

[0128] in, This represents the overall objective function value. This represents all learnable parameters of the model. Indicates the first training subset The true values ​​of soil nutrients for each sample. This represents the number of leaf nodes in the tree. Represents the fraction of the leaf node. , Both represent hyperparameters that control the complexity of the model.

[0129] In order to achieve the best prediction results, considering that the influencing factors and mechanisms of different soil nutrient contents may be different, this invention selects four nutrients, namely organic matter, nitrogen, phosphorus and potassium, to establish independent prediction models. That is, organic matter, nitrogen, phosphorus and potassium are used as output quantities separately during model training.

[0130] Simultaneously, hyperparameter optimization is also required during model training, specifically:

[0131] A training subset is used as input to the model. The multi-source remote sensing features (including spectral features, thermal infrared features, and microenvironment geometric features) in this subset serve as the model's independent variables, while the corresponding laboratory measurements of four nutrients serve as the model's dependent variables. Based on the selected model framework, automated hyperparameter tuning methods such as grid search or Bayesian optimization are employed to optimize key hyperparameters on the training subset using K-fold cross-validation (e.g., 5-fold or 10-fold). For the random forest model, the hyperparameters to be optimized mainly include the number of decision trees, the maximum tree depth, and the minimum number of samples required for internal node splits. For the gradient boosting decision tree model, the hyperparameters to be optimized mainly include the learning rate, the maximum tree depth, the number of boosting iterations, and the subsampling ratio. The optimization process aims to maximize the average coefficient of determination or minimize the average root mean square error on the validation subset.

[0132] Then, the model is validated and its performance is evaluated: specifically as follows:

[0133] The final model after hyperparameter optimization was rigorously evaluated using an independent validation subset that was not involved in model building and parameter optimization during training. The evaluation used the following quantitative metrics:

[0134] Coefficient of determination :

[0135]

[0136] in, This represents the average of the actual soil nutrient values. The number of samples for the independent validation subset; and The closer a value is to 1, the stronger the model's ability to explain data variation.

[0137] Root mean square error :

[0138]

[0139] Among them, RMSE reflects the absolute error between the model's predicted value and the actual value. The smaller the value, the higher the accuracy of the model's prediction.

[0140] Finally, based on the comprehensive performance evaluation results obtained on the independent validation subset, the best-performing model for each of the multiple candidate models trained for the four soil nutrients was selected and solidified into the final soil nutrient prediction model corresponding to that nutrient, so as to be applied to the subsequent step of mapping the spatial distribution of soil nutrients in the entire photovoltaic power station under test.

[0141] The system acquires multi-source remote sensing features of the photovoltaic power station under test, inputs them into a trained machine learning model, and generates soil nutrient results for the photovoltaic power station under test.

[0142] In this embodiment, multi-source remote sensing features of the photovoltaic power station under test are acquired and pre-stored into the corresponding trained machine learning model to obtain the soil nutrient results of the photovoltaic power station under test, thereby generating a spatial distribution map of soil nutrients for the entire photovoltaic power station under test. Specifically:

[0143] For each pixel within the effective analysis area mask of the photovoltaic power station to be tested, multi-source remote sensing feature vectors are systematically extracted and nutrient prediction is performed; for road areas outside the mask, no prediction is made and they are left blank or specially marked in the result map.

[0144] (1) Standardize the extraction of multi-source remote sensing features across the entire area; for each pixel in the digital orthophoto, digital surface model and microenvironment quantification results within the target photovoltaic power station area, systematically extract multi-source remote sensing feature vectors that are completely consistent with the model training stage; specifically, for each pixel in the raster data layer, extract the reflectance values ​​of blue light, green light, red light, red edge and near-infrared bands from the orthophoto based on its geographic coordinates, and calculate spectral indices such as soil organic carbon index, normalized difference vegetation index, normalized difference vegetation index based on green light band, enhanced vegetation index, soil-regulated vegetation index, and red edge chlorophyll index according to the established formula; extract the surface temperature value of the corresponding location from the surface temperature distribution map; extract the symbolic horizontal distance value and microenvironment type classification label of the pixel from the microenvironment quantification results. Thus, construct a feature vector for each pixel within the power station area that is completely consistent with the feature dimensions and physical meaning of the training subset;

[0145] (2) Perform spatial prediction of soil nutrients based on machine learning; take the complete multi-source remote sensing feature vector constructed for each pixel in the above steps as input data and input it into the four independent soil nutrient prediction models that have been trained and optimized above—namely, the organic carbon content prediction model, the total nitrogen content prediction model, the available phosphorus content prediction model, and the available potassium content prediction model. Each model performs parallel calculations on the input feature vector and outputs the predicted values ​​of soil organic carbon content, total nitrogen content, available phosphorus content, and available potassium content corresponding to that pixel; this process is executed cyclically on all pixels in the entire area through automated scripts or parallel computing technology until each pixel in the power station area obtains the quantitative prediction results of its four key soil nutrients;

[0146] (3) Complete the synthesis and output of soil nutrient spatial distribution maps. The four nutrient prediction results corresponding to each geographic pixel obtained in the above steps are reorganized and rendered according to their spatial coordinates; using the geographic information system platform, the discrete pixel prediction values ​​are converted into continuous raster image data, thereby generating four thematic maps of soil nutrient spatial distribution with the same spatial resolution and geographic reference as the original remote sensing data, namely: soil organic carbon content distribution map, soil total nitrogen content distribution map, soil available phosphorus content distribution map and soil available potassium content distribution map; these distribution maps can be visualized using gradient color bands to intuitively show the spatial heterogeneity of each nutrient in different microenvironments such as under the photovoltaic panel, at the front edge of the panel and between the panels; the final generated thematic maps are output in standard geospatial data formats such as GeoTIFF, which can be directly applied to practical scenarios such as precise variable fertilization decision support, power plant ecological health status assessment and long-term dynamic monitoring and management of soil nutrients.

[0147] Specific embodiments have been used to illustrate the principles and implementation methods of this invention. The descriptions of the embodiments above are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.

[0148] Those skilled in the art will recognize that the embodiments described herein are intended to help the reader understand the principles of the invention, and should be understood that the scope of protection of the invention is not limited to such specific statements and embodiments. Those skilled in the art can make various other specific modifications and combinations based on the technical teachings disclosed in this invention without departing from the spirit of the invention, and these modifications and combinations are still within the scope of protection of this invention.

Claims

1. A method for predicting soil nutrients in photovoltaic power plants based on multi-source remote sensing data fusion, characterized in that, Includes the following steps: Multispectral, thermal infrared, and 3D remote sensing data of the target photovoltaic power station are acquired and preprocessed to generate orthophotos, surface temperature distribution maps, and digital surface models. Based on the orthophotos and digital surface models, road areas are identified and extracted, generating a mask for effective analysis areas excluding roads. Within the mask, the projection areas of photovoltaic panels and their front edges are extracted. By calculating symbolic horizontal distances, three types of microenvironments are defined: the area under the panels, the area at the front edge of the panels, and the area between the panels. Specifically, within the mask, a planar segmentation method is used to identify the 3D point cloud of the photovoltaic panels from the digital surface model, projecting it vertically onto a horizontal plane. By calculating the 2D convex hull, several photovoltaic panel projection areas are generated. The average elevation of each side of each photovoltaic panel projection area is compared, and the edge with the lowest elevation is taken as the front edge of the panel. Each pixel within the mask is defined as... Then like a pixel The horizontal projection coordinates are Judgment point If a photovoltaic panel is located within the projection area of ​​any photovoltaic panel, then the photovoltaic panel corresponding to that projection area is selected as the dominant photovoltaic panel; otherwise, the panel whose front edge is located at a certain point is selected. All photovoltaic panels on the north side are considered as candidate photovoltaic panels, and calculation points are calculated. The shortest vertical distance to the front edge of all candidate photovoltaic panels is used to select the candidate photovoltaic panel with the smallest distance from its front edge, which is then selected as the dominant photovoltaic panel; calculation points The shortest Euclidean distance to the boundary of the dominant photovoltaic panel projection area is used as the absolute value of the symbolic horizontal distance to determine the point. If a cell is located inside the projection area of ​​the dominant photovoltaic panel, the symbolic horizontal distance is negative; otherwise, it is positive. All pixels with a symbolic horizontal distance less than 0 are classified as the under-panel region, and a classification label is defined for this region. All pixels with a symbolic horizontal distance greater than or equal to 0 and less than or equal to a preset distance threshold are classified as the panel-front region, and a classification label is defined for this region. All pixels with a symbolic horizontal distance greater than a preset distance threshold are classified as the inter-panel region, and a classification label is defined for this region. Soil sampling points are set up in three types of microenvironments to collect soil samples and determine their nutrient content, which is used as the true soil nutrient data. Simultaneously, multi-source remote sensing features are generated. Using the multi-source remote sensing features as input and the true soil nutrient data as output, a machine learning model is trained to generate a trained machine learning model. The multi-source remote sensing features of the photovoltaic power station under test are obtained, input into the trained machine learning model, and the soil nutrient results of the photovoltaic power station under test are generated.

2. The method for predicting soil nutrients in photovoltaic power plants based on multi-source remote sensing data fusion according to claim 1, characterized in that, The process of acquiring multispectral, thermal infrared, and 3D remote sensing data of a target photovoltaic power station, performing data preprocessing, and generating orthophotos, surface temperature distribution maps, and digital surface models is as follows: A high-precision positioning UAV equipped with a multispectral camera, a thermal infrared camera, and 3D information acquisition equipment is configured to perform a flight mission according to preset flight parameters, acquiring multispectral, thermal infrared, and 3D remote sensing data of the target photovoltaic power station to obtain multi-source remote sensing data; radiometric calibration, atmospheric correction, geometric correction, and orthorectification are performed on the multispectral data to generate an orthophoto of the surface spectral reflectance; radiometric calibration, atmospheric correction, and geometric correction are performed on the thermal infrared data to generate a surface temperature distribution map containing surface brightness and temperature information; it is determined whether the 3D remote sensing data is from optical photogrammetry. If so, a digital surface model is generated through feature point matching and interpolation; otherwise, if it is from lidar measurement, filtering, classification, and interpolation are performed to generate a digital surface model.

3. The method for predicting soil nutrients in photovoltaic power plants based on multi-source remote sensing data fusion according to claim 1, characterized in that, The process of identifying and extracting road regions based on orthophotos and digital surface models, and generating an effective analysis region mask excluding roads, is as follows: Based on orthophotos and digital surface models, road regions are identified and vectorized using image segmentation and classification methods to generate a vector polygon layer of roads inside the target photovoltaic power station; the vector polygon layer of roads inside the target photovoltaic power station is used as the non-analysis region, and an effective analysis region mask excluding roads is generated through spatial analysis.

4. The method for predicting soil nutrients in photovoltaic power plants based on multi-source remote sensing data fusion according to claim 1, characterized in that, The process of setting up soil sampling points in three types of microenvironments, collecting soil samples, and measuring their nutrient content to serve as true soil nutrient data, while simultaneously generating multi-source remote sensing features, is as follows: A stratified random sampling method is used to set up soil sampling points in the three types of microenvironments, and the coordinates of each soil sampling point are recorded; soil samples are extracted from each soil sampling point to obtain the soil organic matter content, nitrogen content, phosphorus content, and potassium content, which are then used as true soil nutrient data; for each soil sampling point, combined with its coordinates, the original surface reflectance values ​​of blue, green, red, red-edge, and near-infrared bands at each soil sampling point's coordinates are extracted from its corresponding orthophoto image; derived spectral indices sensitive to soil and vegetation attributes are calculated, including soil organic carbon index, normalized difference vegetation index, normalized difference vegetation index based on green band, enhanced vegetation index, soil-regulated vegetation index, and red-edge chlorophyll index, which are then used as spectral features. For each soil sampling point, the surface temperature value at the coordinates of each soil sampling point is extracted from the corresponding surface temperature distribution map and used as a thermal infrared feature. At the same time, the symbolic horizontal distance value and the classification labels corresponding to the three types of microenvironments are used as the microenvironment geometric features. Finally, the spectral features, thermal infrared features, and microenvironment geometric features are used as multi-source remote sensing features.

5. The method for predicting soil nutrients in photovoltaic power plants based on multi-source remote sensing data fusion according to claim 4, characterized in that, The formulas for calculating the soil organic carbon index, normalized difference vegetation index, normalized difference vegetation index based on green light band, enhanced vegetation index, soil-regulated vegetation index, and red-edged chlorophyll index are as follows: in, Indicates the soil organic carbon index. Indicates the reflectivity of the blue light band. Indicates the reflectivity in the red light band. Indicates the reflectivity in the green light band. Indicates the normalized difference vegetation index. Indicates near-infrared reflectivity. This represents the normalized difference vegetation index based on the green light band. Indicates an enhanced vegetation index. Indicates the soil-modifying vegetation index. Indicates the soil conditioning coefficient. Indicates the red-edged chlorophyll index. This indicates the reflectivity of the red-edge band.

6. The method for predicting soil nutrients in photovoltaic power plants based on multi-source remote sensing data fusion according to claim 1, characterized in that, The machine learning model is either a random forest model or a gradient boosting decision tree model.

7. The method for predicting soil nutrients in photovoltaic power plants based on multi-source remote sensing data fusion according to claim 1, characterized in that, The hyperparameters of a random forest model include the number of decision trees, the maximum depth of the trees, and the minimum number of samples required for internal node splits.

8. The method for predicting soil nutrients in photovoltaic power plants based on multi-source remote sensing data fusion according to claim 1, characterized in that, The hyperparameters of the gradient boosting decision tree model include the learning rate, the maximum depth of the tree, the number of boosting iterations, and the subsampling ratio.

9. The method for predicting soil nutrients in photovoltaic power plants based on multi-source remote sensing data fusion according to claim 1, characterized in that, Using multi-source remote sensing features as input and soil nutrient ground truth data as output, when training the machine learning model, any one of the soil organic matter content, nitrogen content, phosphorus content, and potassium content in the soil nutrient ground truth data is used as an independent output to train the machine learning model.

Citation Information

Patent Citations

  • Soil heavy metal content identification method based on unmanned aerial vehicle image and machine learning

    CN117740694A

  • Intelligent measuring and calculating method and system for forest carbon reserve

    CN120543029A