A soil microbial community hyperspectral remote sensing monitoring method and storage medium

By generating pseudo-hyperspectral data through generative adversarial networks and semi-supervised learning, and combining physical information neural networks and nitrogen cycle kinetic equations, the problems of difficulty in obtaining labeled samples and profile information in hyperspectral monitoring of soil microbial communities are solved, and large-scale rapid and accurate soil microbial community monitoring and nitrogen cycle parameter inversion are achieved, reducing costs and improving the reliability of results.

CN120408098BActive Publication Date: 2025-09-30XIAN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510884714.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-09-30
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

Existing hyperspectral monitoring technology for soil microbial communities has problems such as difficulty in obtaining labeled samples, difficulty in obtaining soil profile information, and difficulty in parameterizing the nitrogen cycle, resulting in high monitoring costs, small scope, and high uncertainty in results, making it difficult to meet the needs of large-scale, rapid, and non-destructive monitoring.

Method used

Generative adversarial networks and semi-supervised learning methods are used to generate pseudo-hyperspectral data. Combined with physical information neural networks and nitrogen cycle kinetic equations, the uncertainty of monitoring results is assessed through data assimilation methods. A semi-supervised learning framework and physical information neural networks are constructed to infer the microbial distribution and hydrothermal status of soil profiles, and to invert key process parameters of the nitrogen cycle.

Benefits of technology

It achieves accurate prediction of soil microbial communities under conditions of sparse labeled samples, reduces monitoring costs and time, improves the accuracy of profile information acquisition, provides precise inversion and uncertainty assessment of nitrogen cycle process parameters, and enhances the reliability of monitoring results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408098B_ABST
    Figure CN120408098B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of soil microbial monitoring, and discloses a soil microbial community hyperspectral remote sensing monitoring method and storage medium. The method comprises: constructing a generative adversarial network based on data enhancement and semi-supervised learning to generate pseudo-hyperspectral data that conforms to physical constraints and predict soil microbial community parameters; constructing a physical information neural network, introducing constraints from the soil water and heat transport physical equation, and inferring the microbial distribution and hydrothermal state at different depths of the soil profile; embedding the nitrogen cycle kinetic equation in the physical information neural network, expanding the network output variables and physical constraint losses, and optimizing the model parameters and process parameters. The present invention uses GAN data enhancement and semi-supervised learning technology to maintain the prediction accuracy of microbial community indicators when the number of labeled samples is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of soil microbial monitoring, and more particularly to a soil microbial community hyperspectral remote sensing monitoring method and a storage medium. Background Art

[0002] With the development of precision agriculture and soil ecosystem protection, monitoring soil microbial communities has become increasingly critical, with important implications for crop yields, soil quality, and environmental protection. Soil microorganisms participate in key processes in the carbon and nitrogen cycles, playing an irreplaceable role in nutrient transformation, environmental pollutant degradation, and soil structure improvement.

[0003] Currently, soil microbial community monitoring primarily relies on laboratory analytical methods, such as DNA extraction and sequencing, culture counting, and functional gene analysis. While highly accurate, these methods suffer from significant sampling destructiveness, high analysis costs, time-consuming processes, and limited spatial coverage, making them difficult to meet the needs of large-scale, rapid, and non-destructive monitoring. Hyperspectral remote sensing technology, by acquiring spectral information from surface reflectance, enables rapid, non-destructive, and large-scale monitoring of soil properties, and has made progress in monitoring physical and chemical properties such as soil organic matter and texture. However, the application of hyperspectral technology to soil microbial community monitoring still faces numerous challenges.

[0004] The existing hyperspectral monitoring technology of soil microbial communities has the following main problems: it is difficult to obtain labeled samples, and the high cost of laboratory analysis leads to sparse labeled training data, which limits the performance of deep learning models; profile information is difficult to obtain, and existing hyperspectral technology mainly targets surface characteristics, making it difficult to directly obtain microbial information at different depths of the soil profile; there is a lack of effective combination of physical processes and data-driven, and it either relies entirely on statistical correlations and ignores physical process constraints, or relies entirely on mechanism models but is difficult to parameterize; key processes of the nitrogen cycle are difficult to parameterize, and it is difficult to quickly obtain microbial-driven nitrogen cycle process parameters on a large scale; the lack of uncertainty assessment of the results reduces the reliability of decision support.

[0005] Therefore, there is an urgent need for a hyperspectral remote sensing monitoring method for soil microbial communities that can overcome the above technical problems, achieve accurate prediction under data-sparse conditions, infer information from the surface to the profile, integrate physics and data-driven, and accurately invert process parameters. Summary of the Invention

[0006] The present invention provides a soil microbial community hyperspectral remote sensing monitoring method and storage medium, which solves the technical problems in related technologies such as difficulty in obtaining labeled samples, difficulty in obtaining soil profile information, and difficulty in parameterizing nitrogen cycles.

[0007] The present invention provides a soil microbial community hyperspectral remote sensing monitoring method, comprising:

[0008] Based on data enhancement and semi-supervised learning of generative adversarial networks, we obtain hyperspectral remote sensing data and labeled hyperspectral microbial data pairs, construct a generative adversarial network to generate pseudo-hyperspectral microbial data pairs, build a semi-supervised learning framework, and output a microbial community parameter prediction model;

[0009] Profile inference based on physical information neural network: Based on the output of the prediction model, combined with hyperspectral features, meteorological data and spatial coordinates, a physical information neural network is constructed and the constraints of soil water and heat transport equations are introduced to infer the microbial distribution and water and heat status at different depths;

[0010] The nitrogen cycle dynamics equation is embedded in the physical information neural network, and the key process parameters of the nitrogen cycle are output using the profile microbial distribution as the driving force;

[0011] The key process parameters of the nitrogen cycle and the profile microbial distribution are combined with the observation data using the data assimilation method. The hyperspectral data of the target area are input to generate monitoring results of the spatial distribution of soil microbial communities, the profile microbial abundance distribution, and the distribution of nitrogen cycle parameters, and uncertainty assessment is performed.

[0012] Furthermore, the step of generating pseudo hyperspectral data that complies with physical constraints by the generative adversarial network includes:

[0013] A generative adversarial network (GAN) consisting of a generator and a discriminator is constructed. The generator receives a random noise vector and a microbial parameter condition vector as input to generate hyperspectral data that conforms to the true distribution.

[0014] The generator and discriminator are trained using an alternating optimization strategy until the generator can produce pseudo-hyperspectral data that conforms to physical constraints.

[0015] By inputting different microbial parameter condition vectors, corresponding hyperspectral data are generated to form a large number of pseudo-labeled samples;

[0016] A semi-supervised learning framework is constructed to use real-distributed hyperspectral data, pseudo-hyperspectral data that conforms to physical constraints, and pseudo-labeled samples for model training.

[0017] Furthermore, the constructed semi-supervised learning framework includes a student model and a teacher model. The student model updates its parameters through supervised training, and the teacher model parameters are the exponential moving average of the student model parameters; the loss function of the semi-supervised learning framework includes supervision loss, consistency loss and regularization loss.

[0018] Furthermore, the construction of the physical information neural network includes:

[0019] Construct a neural network consisting of a forward propagation part and a physical constraint part;

[0020] To ensure that the output of the neural network satisfies the physical laws, the physical equation constraints of the soil water and heat transport process are introduced;

[0021] Construct a total loss function that includes data fitting loss and physical constraint loss;

[0022] Generative adversarial networks are used to generate pseudo-section data that meets physical constraints, thereby enhancing the training effect of physical information neural networks.

[0023] Furthermore, the physical equation constraints include a moisture migration equation and a heat conduction equation.

[0024] Furthermore, the step of embedding the nitrogen cycle kinetic equation in the physical information neural network includes:

[0025] Add descriptive equations for key processes of the nitrogen cycle to the physical information neural network framework;

[0026] Expand the output variables of the physical information neural network to add predictions of ammonium nitrogen and nitrate nitrogen concentrations, as well as outputs of nitrification and denitrification rates that vary with depth;

[0027] Expand the physical constraint loss of the physical information neural network and add the residual term of the nitrogen cycle equation;

[0028] Add parameter value range constraints to ensure that the parameters obtained by inversion are physically reasonable.

[0029] Furthermore, the descriptive equations for the key processes of the nitrogen cycle include the nitrification process equation and the denitrification process equation, which describe the dynamic changes of ammonium nitrogen and nitrate nitrogen, as well as the effects of microbial abundance, water content, temperature and organic carbon on the nitrification and denitrification processes.

[0030] Furthermore, the step of utilizing the data assimilation method includes adopting an ensemble Kalman filter method to update parameter estimates when observation data arrives by constructing multiple ensemble members of model parameters.

[0031] Furthermore, the uncertainty assessment includes:

[0032] Utilize the ensemble prediction method to calculate the prediction interval of each monitoring result;

[0033] Generate uncertainty maps to identify areas of low forecast reliability;

[0034] Provide confidence assessment of prediction results and provide reliability reference for decision-making applications;

[0035] Output quality control report of monitoring results, including model performance indicators and applicability evaluation.

[0036] The present invention provides a storage medium comprising a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, they are used to implement the above-mentioned soil microbial community hyperspectral remote sensing monitoring method.

[0037] The beneficial effects of the present invention are: through GAN data augmentation and semi-supervised learning technology, the prediction accuracy of microbial community indicators can be maintained when the number of labeled samples is reduced, and even slightly improved in the prediction of certain functional groups, which reduces the dependence on field sampling and laboratory analysis, and reduces monitoring costs and time;

[0038] Based on physical information neural networks, the microbial distribution and hydrothermal state of soil profiles can be inferred from surface hyperspectral remote sensing data. The average relative error of profile microbial abundance prediction is reduced, and the average relative error of profile hydrothermal state prediction is reduced. This enables the "one-look-at-a-glance" capability from the surface to the profile, reducing the workload of traditional drilling sampling.

[0039] The inversion of key process parameters of the soil microbial-driven nitrogen cycle has been achieved. These parameters are crucial to nitrogen cycle models but difficult to measure directly using traditional methods. This provides core parameter support for precise nitrogen fertilizer management and environmental impact assessment, making large-scale nitrogen transformation monitoring possible.

[0040] By integrating physical process constraints and data-driven technology, it can not only provide accurate prediction results, but also quantify the uncertainty of the prediction, provide reliability assessment for decision-making, and enhance the application value of monitoring results in precision agriculture and environmental management. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 This is a flow chart of a soil microbial community hyperspectral remote sensing monitoring method of the present invention;

[0042] Figure 2 is a flow chart of step 1 in the present invention;

[0043] Figure 3 It is a flow chart of step 2 in the present invention;

[0044] Figure 4 It is a flow chart of step 3 in the present invention;

[0045] Figure 5 is a flow chart of step 4 in the present invention;

[0046] Figure 6 This is a comparative diagram of the microbial community prediction accuracy and data efficiency of different methods in the present invention;

[0047] Figure 7It is a thermal diagram of the spatial distribution and uncertainty assessment of key parameters of the nitrogen cycle in the present invention;

[0048] Figure 8 It is a line graph of parameter prediction performance and verification at different depths of the soil profile in the present invention. DETAILED DESCRIPTION

[0049] The subject matter described herein will now be discussed with reference to example embodiments. It should be understood that these embodiments are discussed solely to enable those skilled in the art to better understand and implement the subject matter described herein, and that the functions and arrangements of the elements discussed may be varied without departing from the scope of this specification. Various examples may omit, substitute, or add various processes or components as needed. Furthermore, features described in some examples may be combined in other examples.

[0050] At least one embodiment of the present invention discloses a method for monitoring soil microbial communities using hyperspectral remote sensing, such as Figures 1 to 5 Shown, including:

[0051] Step 1: Data augmentation and semi-supervised learning based on generative adversarial networks are used to obtain hyperspectral remote sensing data and labeled hyperspectral microbial data pairs. A generative adversarial network is constructed to generate pseudo-hyperspectral microbial data pairs, and a semi-supervised learning framework is constructed to output a microbial community parameter prediction model.

[0052] This step uses Generative Adversarial Networks (GANs) to address the problem of sparse labeled samples in hyperspectral remote sensing monitoring and improve the generalization ability of deep learning models. It should be understood that the specific implementation is as follows:

[0053] Step 1.1, data preparation;

[0054] Obtain a small amount of labeled hyperspectral microbial data pairs (labeled data) and a large amount of unlabeled hyperspectral data (unlabeled data). Hyperspectral features in the labeled data include surface reflectance curves (typically covering the 400 to 2500 nm band), and microbial labels include community composition (such as phylum-level taxonomic abundance), diversity indices (such as the Shannon index), or the abundance of specific functional groups.

[0055] Step 1.2, GAN model construction;

[0056] Construct a GAN model suitable for the characteristics of hyperspectral data, including two parts: generator and discriminator. Receive random noise vector and conditional vector (microbial parameters) as input to generate hyperspectral data that conforms to the real distribution; discriminator Distinguish between real hyperspectral data and generated pseudo hyperspectral data.

[0057] According to an embodiment of the present application, the specific implementation of the GAN model includes the following structure:

[0058] The generator adopts an encoder-decoder structure, where the encoder transforms the random noise vector and conditional vector Encoded into latent features, the decoder reconstructs the latent features into hyperspectral data:

[0059] The encoder consists of three convolutional layers, each followed by batch normalization and LeakyReLU activation function; the decoder consists of three transposed convolutional layers for upsampling. The last layer uses the Tanh activation function to ensure that the output value range is between [-1, 1], and then is mapped to the interval [0, 1] as the reflectivity value through linear transformation;

[0060] The discriminator adopts a multi-layer convolutional neural network structure, which includes 4 convolutional layers. Each layer is followed by batch normalization and LeakyReLU activation function. The last layer outputs a single scalar, which indicates the probability that the input data is a real sample.

[0061] Optionally, in some embodiments, the generator may adopt an encoder-decoder structure based on an attention mechanism, wherein a self-attention layer is introduced between the encoder and the decoder to enhance the ability to capture long-distance dependencies between features, which is particularly suitable for modeling complex correlations between different bands in hyperspectral data.

[0062] In other embodiments, the discriminator may adopt an objective function based on the Wasserstein distance, and ensure training stability through a gradient penalty term. The loss function is modified as follows:

[0063] ;

[0064] in Represents the loss function of the discriminator (DiscriminatorLoss); Represents the real data distribution Sampling , discriminator right The expected value of the output, is a real hyperspectral sample, is the probability distribution of real hyperspectral data; Represents the noise distribution Sampling , through the generator Generate fake samples Post-discriminator The expected value of the output, is a random noise vector, is the noise distribution, For the generated hyperspectral data; Represents the weight coefficient of the gradient penalty term, which controls the contribution of the gradient penalty term in the loss function; Represents a uniform interpolation distribution between real samples and generated samples Sampling , discriminator The expectation of the square of the difference between the L2 norm of the gradient of and 1, express right The gradient, represents the L2 norm; Representation Discriminator Input Output.

[0065] ;

[0066] in Represents the loss function of the generator (GeneratorLoss); Represents the noise distribution Sampling , generator Generate fake samples Post-discriminator The expected value of the output; represents a random noise vector, is the noise distribution; Representation Discriminator For the generator Output of the generated samples.

[0067] The loss functions of the generator G and the discriminator D are:

[0068] ;

[0069] in represents the generator loss; Represents the noise distribution Sampled random noise vector , is the noise distribution; Represents the distribution of microbial parameters Sampled conditional vector , is the probability distribution of microbial parameters; Representation Discriminator Generate data and conditions Output; Representation Generator by and Pseudo-hyperspectral data generated for input; Represents the expectation operator;

[0070] ;

[0071] in represents the discriminator loss; Represents the distribution of real hyperspectral data Sampling , is the probability distribution of real hyperspectral data; Represents the distribution of microbial parameters Sampled conditional vector ; Representation Discriminator For real samples and conditions The logarithmic loss of the output; Represents the discriminator's response to input data In the conditions Below is the probability estimate of the real sample; Represents the noise distribution Sampled random noise vector ; Representation Discriminator To generate samples and conditions The logarithmic loss of the output; Represents the expectation operator.

[0072] In the application scenario of farmland soil microbial monitoring, this GAN model can be used to generate hyperspectral data under different conditions of nitrogen cycle functional bacterial abundance. For example, in the application of precision fertilization decision support, by inputting different nitrifying bacteria abundance conditions, corresponding hyperspectral data is generated, which can compensate for the sparse field sampling points and provide data support for subsequent fertilization zoning.

[0073] Step 1.3, GAN training and data generation;

[0074] Training the generator using an alternating optimization strategy and the discriminator , until the generator can generate pseudo hyperspectral data that meets the physical constraints. After the training is completed, by inputting different microbial parameter condition vectors , generate corresponding hyperspectral data, and form a large number of pseudo-labeled samples. The newly generated pseudo samples must also meet physical constraints, such as the reflectance value range is The spectral curve is continuous and smooth.

[0075] Step 1.4, semi-supervised learning model construction;

[0076] A semi-supervised learning framework is constructed based on real labeled samples, generated pseudo-labeled samples, and a large number of unlabeled samples. According to one embodiment of the present application, the semi-supervised learning framework adopts a MeanTeacher model structure, which includes a student model and a teacher model. The student model updates its parameters through supervised training, and the teacher model parameters are the exponential moving average of the student model parameters.

[0077] The specific implementation of this semi-supervised learning framework includes:

[0078] The student model and the teacher model use the same network structure, which consists of four convolutional blocks and two fully connected layers. Each convolutional block contains a convolutional layer, a batch normalization layer, and a ReLU activation function.

[0079] For labeled data (including real labeled samples and high-quality pseudo-labeled samples), the mean squared error (MSE) between the predicted value and the true label is calculated as the supervision loss ;

[0080] For unlabeled data, different data enhancements are applied to the same input, and predictions are made by the student model and the teacher model respectively. The mean square error of the prediction results of the two is calculated as the consistency loss. ; Regularization loss L2 regularization is used to prevent the model from overfitting.

[0081] Optionally, in some embodiments, the semi-supervised learning framework can employ a pseudo-labeling approach. This involves first training an initial model using labeled data, then using this model to predict unlabeled data to generate pseudo-labels. Finally, both the labeled and pseudo-labeled data are used for model training. To control the quality of the pseudo-labels, a confidence threshold can be set to ensure that only high-confidence pseudo-labels are used.

[0082] In other implementations, the semi-supervised learning framework can employ a mixed consistency regularization (MixMatch) approach to generate more diverse training samples by mixing augmented versions of labeled and unlabeled data. Specifically, multiple augmented versions are generated for each unlabeled sample, and the predictions of these augmented versions are averaged as pseudo-labels. The labeled and unlabeled samples are then mixed to generate a mixed sample for use in computing the loss function.

[0083] The framework contains three loss functions: supervision loss (calculated using real labeled samples and pseudo labeled samples), consistency loss (calculated using unlabeled data) and regularization loss The total loss function is:

[0084] ;

[0085] in represents the total loss function; represents the supervision loss, which measures the prediction error of the model for labeled data; represents the consistency loss, which measures the consistency of the prediction results of the student model and the teacher model on the unlabeled data; Represents regularization loss, usually L2 regularization, to prevent overfitting; Represents consistency loss The weight coefficient of controls its contribution to the total loss; represents the regularization loss The weight coefficient of , which controls its contribution to the total loss.

[0086] In regional soil monitoring scenarios, this semi-supervised learning framework can be used to predict soil microbial functional groups at the regional scale. Application examples include monitoring large agricultural regions, where only a small number of sampling points with laboratory data are needed. By combining these with a large number of locations with only hyperspectral data, this framework can generate a microbial functional map for the entire region, providing decision support for regional soil health assessment and precision fertilization management.

[0087] Through the above steps, a semi-supervised learning model was obtained that can accurately predict soil microbial community parameters from hyperspectral data. Therefore, the model has good generalization ability.

[0088] Step 2: Profile inference based on physical information neural network. Based on the output of the prediction model, combined with hyperspectral features, meteorological data, and spatial coordinates, a physical information neural network is constructed and constraints are introduced into the soil water and heat transport equation to infer the microbial distribution and water and heat status at different depths.

[0089] This step uses the Physics-Informed Neural Networks (PINN) method to solve the problem of inferring the distribution of soil microorganisms and their coupling relationship with hydrothermal processes from surface hyperspectral data. It should be noted that the specific implementation is as follows:

[0090] Step 2.1, PINN network architecture construction;

[0091] Construct a physical information neural network, which includes a forward propagation part and a physical constraint part. The forward propagation part is a multi-layer neural network with the input being the surface hyperspectral feature vector (Extracted from the hyperspectral data preprocessed in step 1), meteorological data vector (Including precipitation ,temperature etc.) and spatial position coordinates and depth coordinates , the output is the microbial abundance at that location and depth , soil moisture content and soil temperature .

[0092] According to one embodiment of the present application, the specific structure of the PINN network includes:

[0093] Input layer: Receives surface hyperspectral feature vectors (reduced to 20 principal components), meteorological data vectors (5-dimensional, including precipitation, temperature, humidity, radiation intensity and wind speed), spatial coordinates (2-dimensional, latitude and longitude or XY coordinates) and depth coordinates (1-dimensional);

[0094] Hidden layer: Contains 5 fully connected layers, each containing 64 neurons, and uses the Swish activation function to improve nonlinear expression capabilities;

[0095] Output layer: Output microbial abundance (can be a multidimensional vector, representing the abundance of different groups), soil moisture content and soil temperature.

[0096] Optionally, in some implementations, the PINN network can adopt a residual network structure and introduce skip connections to alleviate the vanishing gradient problem in deep network training and improve model convergence speed and performance. Specifically, a skip connection is added between every two hidden layers, directly adding the output of the previous layer to the output of the next layer.

[0097] In other implementations, the PINN network can employ a multi-scale fusion architecture, processing input features at different scales through multiple parallel sub-networks, and then fusing these multi-scale features for prediction. For example, three parallel sub-networks can be constructed to process hyperspectral features, meteorological data, and spatial-depth coordinates, respectively, and then the outputs of the three sub-networks can be concatenated or weighted fused. This architecture is particularly suitable for fusing heterogeneous data from multiple sources.

[0098] The network structure can be expressed as:

[0099] ;

[0100] in Indicates the location ,depth ,time The abundance of microorganisms; Indicates the location ,depth ,time Volumetric water content of soil; Indicates the location ,depth ,time soil temperature; Represents a neural network function, which represents the mapping from input features to output variables; 、 、 、 Represent the surface hyperspectral feature vector, meteorological data vector, spatial coordinates and depth coordinates respectively; Represents the set of neural network parameters, including all trainable weights and biases.

[0101] Step 2.2, physical constraint definition;

[0102] To ensure that the output of the neural network satisfies the physical laws, the following physical equation constraints are introduced for the soil water and heat transport process:

[0103] Water transport equation (Richards equation):

[0104] ;

[0105] in Indicates soil volume water content About time The partial derivative of represents the rate of change of water content with time; represents the gradient operator, which represents the spatial derivative; represents the unsaturated hydraulic conductivity, which depends on the water content function; represents the spatial gradient of water potential; Indicates source and sink items, such as root water absorption, indicating water loss or replenishment;

[0106] Heat conduction equation:

[0107] ;

[0108] in It represents the heat capacity of soil, the heat capacity of soil per unit volume; Indicates soil temperature About time The partial derivative of It represents the soil thermal conductivity coefficient, which indicates the heat transfer capacity of the soil; represents the spatial gradient of soil temperature; Indicates the density of water; represents the specific heat capacity of water; represents the water velocity vector, which indicates the flow rate of water in the soil; represents the gradient operator, which represents the spatial derivative;

[0109] Step 2.3, PINN loss function construction;

[0110] The total loss function of PINN consists of two parts: data fitting loss and physical constraint loss:

[0111] ;

[0112] in Represents the total loss function of PINN; Represents the weight coefficient of data fitting loss, controlling contribution to total losses; Represents data fitting loss, which measures the error between model output and observed data; Represents the weight coefficient of physical constraint loss, controlling contribution to total losses; Represents the physical constraint loss, which measures the degree to which the model output satisfies the physical equations;

[0113] Data fitting loss Calculation based on profile observation data from a small number of borehole sampling points:

[0114] ;

[0115] in represents the total number of observation points; Represents data fitting loss, which measures the error between model output and observed data; Indicates the Measured values ​​of microbial abundance at each observation point; Indicates the The model-predicted value of microbial abundance at each observation point; Indicates the The measured value of soil volume water content at each observation point; Indicates the The model predicted value of soil volumetric water content at each observation point; Indicates the Measured soil temperature values ​​at each observation point; Indicates the The soil temperature model prediction value of each observation point.

[0116] Physical constraint loss Based on the residual calculation of the above physical equation, the automatic differentiation technique is used to evaluate the difference between the left and right sides of the equation:

[0117]

[0118] in represents the physical constraint loss, Indicates the number of sampling points used to calculate physical constraints (can be virtual points); Represents the Richards equation at point The residual, 、 、 Respectively represent The spatial coordinates, depth coordinates and time coordinates of each observation point; The heat conduction equation is expressed at the point The residual, 、 、 Respectively represent The spatial coordinates, depth coordinates and time coordinates of each observation point.

[0119] Step 2.4, GAN-assisted physical constraint sample generation;

[0120] Using the GAN model trained in step 1, pseudo-profile data that meets physical constraints is generated to enhance the PINN training effect. These pseudo-samples not only conform to the statistical relationship between hyperspectral and microbial abundance, but also meet the constraints of the physical process of water and heat transport.

[0121] Through the above steps, a PINN model was trained to infer the microbial distribution and hydrothermal state at different depths in the soil profile from surface hyperspectral data. In addition, the model not only fits the actual observation data, but also conforms to the physical laws of soil water and heat migration.

[0122] Step 3: embed the nitrogen cycle kinetic equation into the physical information neural network, use the profile microbial distribution as the driving force, and output the key process parameters of the nitrogen cycle;

[0123] This step embeds the nitrogen cycle kinetic equations based on the PINN framework to achieve the inversion of key process parameters of the microbial-driven nitrogen cycle, solving the problem of difficulty in directly obtaining parameters. It should be noted that the specific implementation is as follows:

[0124] Step 3.1, embedding the kinetic equation of the nitrogen cycle process;

[0125] In the PINN framework of step 2, add equations describing the key processes of the nitrogen cycle. Taking nitrification and denitrification as an example, introduce the following equations:

[0126] Nitrification process equation:

[0127] ;

[0128] in represents the spatial gradient of ammonium nitrogen concentration; The concentration of ammonium nitrogen versus time The partial derivative of represents the gradient operator, which represents the spatial derivative; represents the diffusion coefficient of ammonium ions, Diffusion capacity in soil; represents the nitrification rate constant, which is the rate parameter of the nitrification reaction; represents the effect function of environmental factors on nitrification, is the microbial abundance, is the volumetric water content of the soil, is soil temperature;

[0129] ;

[0130] in represents the spatial gradient of nitrate nitrogen concentration; The nitrate nitrogen concentration versus time The partial derivative of represents the diffusion coefficient of nitrate ions, Diffusion capacity in soil; represents the denitrification rate constant, which is the rate parameter of the denitrification reaction; represents the impact function of environmental factors on denitrification, is the microbial abundance, is the volumetric water content of the soil, is the soil temperature, is the organic carbon content; represents the gradient operator, which represents the spatial derivative; represents the nitrification rate constant, which is the rate parameter of the nitrification reaction; Indicates the concentration of ammonium nitrogen in the soil content; represents the effect function of environmental factors on nitrification, is the microbial abundance, is the volumetric water content of the soil, is soil temperature;

[0131] ;

[0132] in represents the effect function of environmental factors on nitrification, is the microbial abundance, is the volumetric water content of the soil, is soil temperature; The function representing the effect of water content on the process rate is shown below; Represents the function of the effect of temperature on the process rate;

[0133] ;

[0134] in represents the impact function of environmental factors on denitrification, is the microbial abundance, is the volumetric water content of the soil, is the soil temperature, is the organic carbon content; Function representing the effect of water content on process rate; Represents the function of the effect of temperature on the process rate; Function representing the effect of organic carbon on process rate.

[0135] According to one embodiment of the present application, these influence functions can be specifically expressed as:

[0136] when hour:

[0137] ;

[0138] otherwise ;

[0139] in Function that represents the effect of water content on process rate; Indicates the volumetric water content of soil; Indicates the lower limit of water content; Indicates the optimum water content; Indicates the upper limit of water content;

[0140] ;

[0141] in Represents the function of the effect of temperature on the process rate; Indicates soil temperature; Indicates the reference temperature; It represents the temperature sensitivity coefficient, which indicates the multiple of reaction rate increase when the temperature rises by 10℃;

[0142] ;

[0143] in Function representing the effect of organic carbon on process rate; Indicates organic carbon content; represents the half-saturation constant, which indicates the saturation effect of organic carbon on the process rate;

[0144] Step 3.2, expand PINN output variables;

[0145] Based on the PINN network output in step 2, the addition of ammonium nitrogen , nitrate nitrogen Concentration prediction and key parameters and Output that varies with depth.

[0146] According to one embodiment of the present application, the extended PINN network structure is based on step 2, and the output layer adds 4 additional output nodes, corresponding to ammonium nitrogen concentration, nitrate nitrogen concentration, nitrification rate constant, and denitrification rate constant, respectively. The network structure can be expressed as:

[0147] ;

[0148] in represents the abundance of microorganisms; Indicates the volumetric water content of soil; Indicates soil temperature; Indicates the concentration of ammonium nitrogen; Indicates the concentration of nitrate nitrogen; represents the nitrification rate constant; represents the denitrification rate constant; Represents the expanded neural network function; 、 、 、 、 They represent the surface hyperspectral feature vector, meteorological data vector, spatial coordinates, depth coordinates and neural network parameters respectively.

[0149] Step 3.3, expand the PINN physical constraint loss;

[0150] Based on the physical constraint loss in step 2, the residual term of the nitrogen cycle equation is added:

[0151] ;

[0152] in represents the loss of physical constraints after expansion; represents the loss of physical constraints; Represents the weight coefficient of the residual of the nitrogen cycle equation, which controls the contribution of the residual of the nitrogen cycle equation to the total loss; Indicates the number of sampling points used to calculate physical constraints; The kinetic equation of ammonium nitrogen is expressed at point The residual of The kinetic equation of nitrate nitrogen is expressed at point The residual of 、 、 Respectively represent The spatial coordinates, depth coordinates and time coordinates of each observation point.

[0153] Step 3.4, parameter constraints are added;

[0154] To ensure that the parameters obtained by inversion are physically reasonable, the parameter value range constraints are added:

[0155] ;

[0156] ;

[0157] in represents the minimum value of the nitrification rate constant; represents the maximum value of the nitrification rate constant; represents the minimum value of the denitrification rate constant; represents the maximum value of the denitrification rate constant; represents the nitrification rate constant output by the model; Represents the denitrification rate constant output by the model.

[0158] These constraints can be implemented by adding penalty terms to the loss function, or by using parameter mapping methods (such as the sigmoid function) to ensure that the output parameters are within a reasonable range.

[0159] Step 3.5, data assimilation and parameter optimization;

[0160] Using data assimilation method, surface hyperspectral data are combined with a small amount of profiles 、 The concentration observation data are combined to optimize the PINN model parameters and process parameters.

[0161] According to one embodiment of the present application, the data assimilation process is implemented using the Ensemble Kalman Filter method, which constructs multiple ensemble members of model parameters and updates parameter estimates when observation data arrives. The specific steps include:

[0162] Initialization: Based on prior knowledge, generate multiple set members (e.g., 100) for PINN model parameters and process parameters;

[0163] Prediction: Run the model using the parameters of each ensemble member to obtain the predicted value of the state variable;

[0164] Update: When new observation data arrives, the parameter value of each ensemble member is updated according to the Kalman gain;

[0165] Iteration: Repeat the prediction and update steps until the parameters converge.

[0166] Optionally, in some implementations, the data assimilation process can employ a particle filter approach, which uses importance sampling and resampling steps to handle nonlinear and non-Gaussian distributions. Specifically, each particle (i.e., a member of a parameter set) is assigned a weight representing the degree to which its predicted value matches the observed value. Resampling is then performed based on the weights to generate a new set of particles, with appropriate perturbations added to maintain particle diversity. This approach is particularly well-suited for strongly nonlinear processes in soil-microbe systems.

[0167] In other embodiments, variational data assimilation methods can be used to optimize model parameters by minimizing a cost function. The cost function typically includes a background error term (deviation from the prior parameter estimate) and an observation error term (deviation from the observed data):

[0168] ;

[0169] in represents the cost function, the objective to be minimized; represents the parameter vector to be estimated; represents the prior parameter estimation vector; represents the background error covariance matrix, which indicates the uncertainty of the prior parameters; represents the transpose operator; Represents the observation operator, and the parameter vector Mapping to observation space; Represents the observation vector, which contains the observation data; represents the observation error covariance matrix, which represents the uncertainty of the observation data; Represents an inverse matrix operation.

[0170] The optimization goal is to minimize the expanded PINN total loss function:

[0171] ;

[0172] in Represents the total loss function of the expanded PINN; Represents the weight coefficient of the extended data loss term; represents the extended data loss term that includes the fitting error of the nitrogen concentration observation data; The weight coefficient representing the extended physical constraint loss; represents the physical constraint loss after expansion.

[0173] In environmental monitoring application scenarios, this method can be used to monitor nitrogen cycle processes in agricultural and natural ecosystems, assess nitrogen loss risks (such as nitrate leaching, nitrogen gas and nitrous oxide emissions), and provide a scientific basis for environmental protection and emission reduction measures.

[0174] Through the above steps, we can obtain the key parameters of soil nitrogen cycle ( and ) model, it can be seen that these parameters are consistent with the physical, chemical and biological process constraints and can explain the actual observed nitrogen concentration data.

[0175] Step 4: Use data assimilation methods to combine key nitrogen cycle process parameters and profile microbial distribution with observational data, input hyperspectral data of the target area, generate monitoring results for the spatial distribution of soil microbial communities, profile microbial abundance distribution, and nitrogen cycle parameter distribution, and perform uncertainty assessment;

[0176] Based on the integrated model trained in the above three steps, hyperspectral remote sensing monitoring of soil microbial communities in the target area is achieved.

[0177] Step 4.1, target area data input;

[0178] Input the hyperspectral remote sensing data of the area to be monitored into the trained model system, including:

[0179] Surface hyperspectral data: hyperspectral images of the target area acquired by satellites, drones, or ground equipment;

[0180] Auxiliary data: meteorological data (temperature, humidity, precipitation, etc.) and geographic location information of the target area;

[0181] Spatial coordinates: geographic coordinate information of each pixel in the target area.

[0182] Step 4.2, model prediction calculation;

[0183] The input data is calculated through the integrated predictive model system:

[0184] First, the semi-supervised learning model trained in step 1 is used to predict surface microbial community parameters from hyperspectral data.

[0185] Then, the physical information neural network trained in step 2 is used to infer the microbial distribution and hydrothermal state at different depths of the soil profile;

[0186] Finally, the key process parameters of the nitrogen cycle driven by microorganisms are obtained through the nitrogen cycle parameter inversion model trained in step 3.

[0187] Step 4.3, remote sensing monitoring results output;

[0188] After model prediction and calculation, complete remote sensing monitoring results of the target area are obtained, including:

[0189] Soil microbial community spatial distribution map: shows the spatial distribution characteristics of different microbial functional groups in the target area;

[0190] Profile microbial abundance distribution: provides information on the abundance distribution of microorganisms at different depths of the soil profile;

[0191] Distribution of nitrogen cycle process parameters: including the spatial distribution of key parameters such as nitrification rate and denitrification rate;

[0192] Soil environmental status information: profile distribution of environmental factors such as soil moisture content and temperature.

[0193] Step 4.4, uncertainty assessment and reliability quantification;

[0194] Perform uncertainty assessment on all forecast results:

[0195] Utilize the ensemble prediction method to calculate the prediction interval of each monitoring result;

[0196] Generate uncertainty maps to identify areas of low forecast reliability;

[0197] Provide confidence assessment of prediction results and provide reliability reference for decision-making applications;

[0198] Output quality control report of monitoring results, including model performance indicators and applicability evaluation.

[0199] Through the above steps, a complete hyperspectral remote sensing monitoring result of soil microbial community was formed, providing a scientific basis for precision agricultural management, environmental monitoring and ecological assessment.

[0200] A storage medium includes a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, it is used to implement the above-mentioned soil microbial community hyperspectral remote sensing monitoring method.

[0201] Here, the present invention provides an implementation example:

[0202] Provided are real application examples of the present invention in monitoring farmland soil microbial communities:

[0203] This application example focuses on hyperspectral remote sensing monitoring of soil microbial communities in a farmland ecosystem. The farmland, covering approximately 100 hectares, primarily grows wheat and corn in rotation. Traditionally, uniform fertilization has resulted in low nitrogen fertilizer utilization efficiency in some areas and the risk of nitrate leaching. This example aims to achieve high-precision monitoring of soil microbial community distribution and key nitrogen cycle parameters through the methods described in this application, providing decision support for precision fertilization.

[0204] In this farmland, a drone-mounted hyperspectral imager (wavelength range 400-2500 nm, spectral resolution 10 nm) was used to acquire surface hyperspectral data. Soil samples were collected at 25 sampling points throughout the farmland. Only 10 of these points were used to determine soil microbial community composition (using 16S rRNA sequencing) and the abundance of nitrogen cycle functional genes. These 10 points served as labeled data; the remaining 15 points only had basic physical and chemical properties measured, serving as unlabeled data. Furthermore, stratified sampling was performed at 5 of the 10 labeled points, with samples taken every 20 cm along the 0-100 cm profile. The microbial composition and nitrogen forms of each layer were determined for use in PINN model training and validation.

[0205] Hyperspectral data preprocessing included atmospheric correction, geometric correction, denoising, and band selection, ultimately selecting 50 characteristic bands for subsequent analysis. Soil microbial data were processed using a bioinformatics analysis pipeline to obtain information on microbial community composition, with a focus on functional bacteria associated with the nitrogen cycle, including nitrifying bacteria (AOB and AOA) and denitrifying bacteria.

[0206] To address the issue of insufficient labeled samples, a conditional GAN ​​model was constructed. The generator employs an encoder-decoder architecture. The input layer receives a 128-dimensional random noise vector and a 10-dimensional microbial parameter vector (including the relative abundance and diversity index of key bacterial groups). This input is processed through three encoding convolutional layers and three decoding transposed convolutional layers, outputting a 50-band hyperspectral feature vector. The discriminator, consisting of four convolutional layers and one fully connected layer, outputs the authenticity discrimination result.

[0207] The model was trained using mini-batch gradient descent with a batch size of 32, an initial learning rate of 0.0002, and the Adam optimizer. During training, the generator was updated every five times the discriminator was trained to prevent premature collapse. To ensure the quality of generated samples, spectral smoothness constraints and inter-band correlation constraints were introduced. After training, 100 new pseudo-hyperspectral microbial data pairs were generated by inputting different microbial parameter conditions.

[0208] When constructing the semi-supervised learning framework, the MeanTeacher structure was adopted. Both the student model and the teacher model adopt a four-layer convolutional neural network structure. For 10 real labeled samples and 100 generated pseudo-labeled samples, the label prediction error is calculated as the supervision loss; for 15 unlabeled samples, different data enhancements (such as random band masks, random noise addition, etc.) are applied, and the student model and teacher model predict the results, and the mean square error of the two predictions is calculated as the consistency loss. During model training, the consistency loss weight coefficient is Set to 0.5 and increase linearly to 1.0 as training progresses; regularization loss weight coefficient Set to 0.001.

[0209] The constructed PINN network consists of five hidden layers, each with 64 neurons, using the Swish activation function. Input features include a 20-dimensional hyperspectral feature vector extracted from the GAN-trained discriminator, a 5-dimensional meteorological data vector (obtained from the nearest weather station), a 2-dimensional spatial coordinate (field grid coordinates), and a 1-dimensional depth coordinate (in the range of 0-100 cm).

[0210] In terms of physical constraints, Richards water transport equation and heat conduction equation are introduced. Parameter settings include: soil unsaturated hydraulic conductivity Using the van Genuchten model parameterization, soil heat capacity and thermal conductivity Estimated based on soil texture and organic matter content. In the loss function, the data fitting loss weight Set to 1.0, the physical constraint loss weight It is initially set to 0.1 and gradually increases to 0.5 during training.

[0211] The training process used data from five profile sampling points, totaling 25 depth-level observations. Furthermore, pseudo-samples generated by a generative adversarial network (GAN) were used to assist in training, enhancing the model's ability to infer profile information under varying soil conditions. After training, the model was able to infer microbial abundance, soil moisture, and temperature distribution at different depths (0-100 cm) from surface hyperspectral data at any location.

[0212] Based on the PINN framework, the nitrogen cycle dynamics equation is embedded. In the nitrification process equation, the conversion of ammonium nitrogen to nitrate nitrogen is affected by microbial abundance, soil moisture and temperature. In implementation, the moisture effect function The parameter settings are: , , ; Temperature influence function middle The value is set to 2.0, the reference temperature is 20℃.

[0213] The expanded PINN network adds four output nodes: ammonium nitrogen concentration, nitrate nitrogen concentration, nitrification rate constant, and denitrification rate constant. To ensure that the parameters are physically reasonable, the nitrification rate constant range is set to 0.01 to 0.5d , the denitrification rate constant ranges from 0.001 to 0.1d The network output is mapped to a reasonable parameter range through the sigmoid function.

[0214] The data assimilation process employed an ensemble Kalman filter to construct 100 ensemble members, each containing different initial PINN and soil parameters. Based on observations of ammonium and nitrate concentrations at five profile sampling points (a total of 25 depth levels), the ensemble member parameters were updated through multiple iterations. Ultimately, the spatial distribution of nitrification and denitrification rates at different locations and depths across the entire farmland was inverted.

[0215] Based on the trained integrated model, a hyperspectral remote sensing monitoring system was conducted on an entire 100-hectare farmland area. The trained model system was fed with hyperspectral image data (2m spatial resolution, 2500 pixels) from the entire farmland acquired by a drone.

[0216] Input data preparation:

[0217] The hyperspectral imagery of the entire farmland was preprocessed, including radiometric correction, geometric registration, and spectral feature extraction, ultimately obtaining a 50-dimensional hyperspectral feature vector for each pixel. Meteorological data for the corresponding period (average daily temperature of 15.2°C, relative humidity of 65%, and precipitation of 12 mm) and the spatial coordinates of each pixel were also collected.

[0218] Model prediction calculation:

[0219] The preprocessed hyperspectral data is predicted through three sub-models in turn:

[0220] First, the surface microbial community parameters were predicted using a semi-supervised learning model to obtain the abundance of nitrifying bacteria, denitrifying bacteria, and microbial diversity index for each pixel.

[0221] Then, soil profile information was inferred using the PINN model to obtain the distribution of microbial abundance, soil moisture content, and temperature at every 20 cm layer within the depth range of 0-100 cm.

[0222] Finally, the nitrification rate constant and denitrification rate constant at each depth level were obtained through the nitrogen cycle parameter inversion model.

[0223] Monitoring result output:

[0224] After model prediction and calculation, we obtained complete remote sensing monitoring results within the farmland, including:

[0225] Spatial distribution map of soil microbial communities: Spatial distribution maps of nitrifying bacteria, denitrifying bacteria, and total microbial diversity were generated, identifying three high nitrification activity areas, two high denitrification activity areas, and four microbial diversity hotspots.

[0226] Profile microbial abundance distribution: Provides a three-dimensional distribution of microbial abundance at five depth levels (0-20cm, 20-40cm, 40-60cm, 60-80cm, 80-100cm), showing that microbial abundance decreases with depth, but local enrichment occurs in the 40-60cm layer;

[0227] Distribution of nitrogen cycle process parameters: The spatial distribution maps of nitrification and denitrification rates were obtained. The nitrification rate varied in the range of 0.05-0.35 d-1-1, and the denitrification rate varied in the range of 0.002-0.08 d-1-1;

[0228] Soil environmental status information: Soil moisture content varies between 22% and 38%, and soil temperature varies between 13.5 and 16.8°C, showing obvious spatial heterogeneity.

[0229] Uncertainty assessment:

[0230] Using an ensemble prediction method, 95% confidence intervals were calculated for each monitoring result. The results showed an average uncertainty of ±12.5% ​​for microbial abundance predictions and ±18.3% for nitrogen cycle parameter predictions. Uncertainty maps were generated, highlighting a 15-hectare area in the southeast corner of the farmland where prediction reliability was low (uncertainty >25%). Additional validation sampling is recommended in this area.

[0231] Based on the inversion results, farmland was divided into five fertilization management zones. In areas with high nitrifying bacteria abundance and rapid nitrification rates, the amount of nitrogen fertilizer applied at one time was reduced, and a split-application strategy was adopted. In areas with strong denitrification, drainage measures were increased to reduce anaerobic conditions, while nitrification inhibitors were used to slow the conversion of ammonium nitrogen to nitrate nitrogen, thereby improving nitrogen fertilizer utilization and reducing nitrogen losses.

[0232] like Figures 6 to 8 As shown, this application example focuses on verifying two main technical effects: improved data utilization efficiency and breakthrough in nitrogen cycle parameterization.

[0233] In this application scenario, compared to traditional methods that require extensive field sampling and laboratory analysis, this method uses only 10 labeled samples and 15 unlabeled samples, reducing data acquisition costs and time. Through GAN data augmentation and semi-supervised learning, the prediction accuracy reaches the level of a traditional model trained with 30 labeled samples.

[0234] Specifically, the root mean square error (RMSE) for nitrifying bacteria abundance prediction is compared as follows: 0.42 for the traditional supervised learning method (30 labeled samples); 0.68 for the traditional supervised learning method (10 labeled samples); and 0.39 for the present invention (10 labeled samples + GAN augmentation + semi-supervised learning). This demonstrates that while reducing the data size by approximately 67%, the present invention not only maintains prediction accuracy but also slightly improves it (RMSE decreases by approximately 7%).

[0235] The present invention also significantly reduces sampling and analysis costs. Traditional methods require microbial community sequencing and functional gene analysis at 30 sites, costing approximately 45,000 yuan and taking about 15 days. However, the present invention only requires microbial analysis at 10 sites, costing approximately 15,000 yuan and taking about 7 days (including model training).

[0236] This invention successfully inverts key soil profile nitrogen cycle parameters from surface hyperspectral data, providing crucial support for precision nitrogen fertilizer management. Traditional methods require obtaining these parameters through indoor incubation experiments or in situ tracer experiments, which are costly, time-consuming, and have limited spatial representation for each measurement point.

[0237] The accuracy of the inversion parameters was verified by comparing them with laboratory measured values. At 25 depth levels of 5 profile verification points, the correlation coefficient between the inverted nitrification rate constant and the indoor culture measured value reached 0.82, with an average relative error of 16.8%; the R 2 The accuracy is 0.78, and the average relative error is 19.3%. This accuracy has good practical value for monitoring soil nitrogen cycle parameters over a large area.

[0238] Adjustments to fertilization management based on inversion parameters improved nitrogen fertilizer utilization efficiency. In a comparative test in a farmland pilot area, the nitrogen fertilizer utilization rate in areas with traditional uniform fertilization was 35.2%, while in areas using the present invention's zoned precision fertilization, it increased to 46.8%, a 33% increase. Furthermore, nitrate leaching monitoring results showed that nitrogen leaching losses in the precision fertilization areas were 28.5% lower than in areas with traditional fertilization, reducing the risk of environmental pollution.

[0239] Furthermore, the uncertainty assessment function of the present invention provides reliability information for decision-making. By quantifying the uncertainty of parameters obtained from the ensemble prediction, conservative fertilization strategies can be recommended for low-reliability areas (uncertainty greater than 30%), further reducing risk.

[0240] The above describes an embodiment of the present invention, but this embodiment is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Ordinary technicians in this field can also make more forms of equivalent embodiments based on the inspiration of this embodiment, all of which are protected by this embodiment.

Claims

1. A soil microbial community hyperspectral remote sensing monitoring method, characterized in that: include: Based on data enhancement and semi-supervised learning of generative adversarial networks, hyperspectral remote sensing data and labeled hyperspectral microbial data pairs are obtained, a generative adversarial network is constructed to generate pseudo hyperspectral microbial data pairs, a semi-supervised learning framework is constructed, and a microbial community parameter prediction model is output. The steps of constructing a generative adversarial network to generate pseudo hyperspectral microbial data pairs include: A generative adversarial network (GAN) consisting of a generator and a discriminator is constructed. The generator receives a random noise vector and a microbial parameter condition vector as input to generate hyperspectral data that conforms to the true distribution. The generator and discriminator are trained using an alternating optimization strategy until the generator can produce pseudo-hyperspectral data that conforms to physical constraints. By inputting different microbial parameter condition vectors, corresponding hyperspectral data are generated to form pseudo-labeled samples; Construct a semi-supervised learning framework to train the model using real-distributed hyperspectral data, pseudo-hyperspectral data that conforms to physical constraints, and pseudo-labeled samples; Based on the output of the prediction model, the physical information neural network is constructed based on profile inference. The hyperspectral characteristics, meteorological data and spatial coordinates are combined to introduce the constraints of the soil water and heat transport equation to infer the microbial distribution and water and heat status at different depths. The construction of the physical information neural network includes: Construct a neural network consisting of a forward propagation part and a physical constraint part; To ensure that the output of the neural network satisfies the physical laws, the physical equation constraints of the soil water and heat transport process are introduced; Construct a total loss function that includes data fitting loss and physical constraint loss; Generate pseudo-profile data that meets physical constraints using a generative adversarial network; Embedding a nitrogen cycle kinetic equation in a physical information neural network, using profile microbial distribution as a driving force, and outputting key process parameters of the nitrogen cycle. The steps of embedding the nitrogen cycle kinetic equation in the physical information neural network include: Add descriptive equations for key processes of the nitrogen cycle to the physical information neural network framework; Expand the output variables of the physical information neural network to add predictions of ammonium nitrogen and nitrate nitrogen concentrations, as well as outputs of nitrification and denitrification rates that vary with depth; Expand the physical constraint loss of the physical information neural network and add the residual term of the nitrogen cycle equation; Add parameter value range constraints to ensure that the parameters obtained by inversion are physically reasonable; The key process parameters of the nitrogen cycle and the profile microbial distribution are combined with the observation data using the data assimilation method. The hyperspectral data of the target area are input to generate monitoring results of the spatial distribution of soil microbial communities, the profile microbial abundance distribution, and the distribution of nitrogen cycle parameters, and uncertainty assessment is performed.

2. The method for monitoring soil microbial communities using hyperspectral remote sensing according to claim 1, wherein: The constructed semi-supervised learning framework includes a student model and a teacher model. The student model updates its parameters through supervised training, and the teacher model parameters are the exponential moving average of the student model parameters. The loss function of the semi-supervised learning framework includes supervision loss, consistency loss, and regularization loss.

3. The method for monitoring soil microbial communities using hyperspectral remote sensing according to claim 1, wherein: The physical equation constraints include a water transport equation and a heat conduction equation.

4. The method for monitoring soil microbial communities using hyperspectral remote sensing according to claim 1, wherein: The descriptive equations for the key processes of the nitrogen cycle include the nitrification process equation and the denitrification process equation, which describe the dynamic changes of ammonium nitrogen and nitrate nitrogen, as well as the effects of microbial abundance, water content, temperature and organic carbon on the nitrification and denitrification processes.

5. The method for hyperspectral remote sensing monitoring of soil microbial communities according to claim 1, characterized in that: The step of using the data assimilation method includes using an ensemble Kalman filter method to update parameter estimates as observation data arrives by constructing multiple ensemble members of model parameters.

6. The method for monitoring soil microbial communities using hyperspectral remote sensing according to claim 1, wherein: The uncertainty assessment includes: Utilize the ensemble prediction method to calculate the prediction interval of each monitoring result; Generate uncertainty maps to identify areas of low forecast reliability; Provide confidence assessment of prediction results and provide reliability reference for decision-making applications; Output quality control report of monitoring results, including model performance indicators and applicability evaluation.

7. A storage medium, characterized in that: The method comprises a memory and one or more processors, wherein the memory stores executable code, and when the one or more processors execute the executable code, the method is used to implement a soil microbial community hyperspectral remote sensing monitoring method according to any one of claims 1 to 6.