Machine learning-based wheat variety soil nitrogen fertilizer gene detection optimization system

By constructing a machine learning-based wheat variety soil fertility nitrogen fertilizer gene detection and optimization system, the problems of insufficient quantification of soil fertility differences and inaccurate nitrogen fertilizer utilization efficiency assessment in traditional wheat variety screening methods have been solved. This system achieves precise coupling between wheat varieties and soil fertility and efficient nitrogen fertilizer utilization, thereby improving the accuracy and efficiency of variety screening and recommendation.

CN122390148APending Publication Date: 2026-07-14INST OF AGRI RESOURCES & REGIONAL PLANNING CHINESE ACADEMY OF AGRI SCI

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INST OF AGRI RESOURCES & REGIONAL PLANNING CHINESE ACADEMY OF AGRI SCI
Filing Date
2026-04-20
Publication Date
2026-07-14

Smart Images

  • Figure CN122390148A_ABST
    Figure CN122390148A_ABST
Patent Text Reader

Abstract

The present application belongs to the field of machine learning, and particularly relates to a wheat variety soil fertility and nitrogen fertilizer gene detection and optimization system based on machine learning. The system comprises a soil fertility basic data acquisition terminal, a basic productivity evaluation server, a regional soil fertility grading device, a variety phenotype monitoring device, a nitrogen fertilizer utilization efficiency analysis platform, a gene feature extraction module, a machine learning optimization master control system and a decision output terminal. By obtaining soil fertility basic data, variety phenotypes and gene feature vectors, the system uses a machine learning model to perform multidimensional data mining and nonlinear correlation modeling, and outputs variety optimization classification suggestions under different soil fertility levels. Through the above system architecture, the present application realizes precise coupling of variety screening and arable land fertility, improves the precision of nitrogen fertilizer utilization efficiency prediction, provides a scientific basis for reducing the amount of chemical fertilizers and increasing efficiency, and can realize optimal allocation of agricultural production resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of machine learning, specifically relating to a machine learning-based wheat variety soil fertility nitrogen fertilizer gene detection and optimization system. Background Technology

[0002] With the cross-disciplinary integration of precision agriculture and modern bio-breeding technology, the selection of superior varieties for wheat, a major grain crop, is gradually shifting from simple morphological screening to a deeper consideration of comprehensive environmental adaptability and resource utilization. As a core element supporting crop yield, the scientific rigor and accuracy of the evaluation system for arable land directly affect the matching effect of superior varieties and cultivation methods. Establishing a comprehensive evaluation model encompassing geographical, soil, and biological characteristics, and quantifying the health status and production potential of arable land, has become an important foundation for improving wheat yield per unit area and achieving green agricultural transformation.

[0003] Wheat variety selection based on arable land productivity grading is a key technological track for optimizing the allocation of planting resources. This technology aims to construct a regionally targeted variety selection mechanism by finely characterizing the soil fertility levels of different ecological zones and combining the phenotypic feedback and nutrient absorption characteristics of wheat under specific fertility conditions, thereby achieving the dual goals of reducing fertilizer use and increasing grain yield.

[0004] However, traditional wheat variety selection methods often focus on general assessments across macro-regions, lacking quantitative boundaries to define cross-regional soil fertility differences. This makes it difficult for selection results to guide precision planting under specific soil fertility conditions. Furthermore, existing assessment index systems are relatively simplistic, failing to deeply couple nitrogen fertilizer agronomical efficiency with basic productivity, and lack efficient feature analysis capabilities when processing multi-source heterogeneous data. Traditional models struggle to integrate complex gene expression information and environmental stress factors, failing to capture the nonlinear relationship between yield fluctuations and resource inputs, thus limiting the predictive accuracy and applicability of variety breeding. Summary of the Invention

[0005] The purpose of this invention is to provide a wheat variety soil fertility nitrogen fertilizer gene detection and optimization system based on machine learning, which can solve the problems in the above-mentioned background technology, such as low matching degree between wheat variety screening and soil fertility, inaccurate evaluation of nitrogen fertilizer utilization efficiency, and lack of cross-regional multi-dimensional data integration mechanism.

[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0007] A machine learning-based wheat variety soil fertility nitrogen fertilizer gene detection and optimization system includes a soil fertility basic data acquisition terminal, a basic productivity assessment server, a regional soil fertility grading device, a variety phenotypic monitoring device, a nitrogen fertilizer use efficiency analysis platform, a gene feature extraction module, a machine learning optimization master control system, and a decision output terminal, wherein:

[0008] The soil fertility basic data acquisition terminal is used to acquire raw yield data of wheat in the target area under the condition of no fertilizer application, soil physicochemical property parameters and historical tillage records, and transmit the above data to the basic productivity assessment server.

[0009] The basic productivity assessment server is used to standardize the received wheat yield data under the no-fertilization treatment and calculate the basic productivity value of the cultivated land.

[0010] The regional soil fertility classification device uses pre-stored national farmland quality bulletin data, combined with the proportion distribution of high-quality farmland, medium-quality farmland and low-quality farmland in a specific region, to compare the obtained basic productivity values ​​with preset soil fertility classification thresholds, and classify wheat fields in each region into high soil fertility level, medium soil fertility level and low soil fertility level.

[0011] The variety phenotypic monitoring device is used to monitor and record the yield data of different wheat varieties under conventional fertilization treatment in real time on sample plots with different soil fertility levels, and to summarize the data to the machine learning optimization master control system.

[0012] The nitrogen fertilizer utilization efficiency analysis platform is used to calculate the nitrogen fertilizer agronomic efficiency based on the yield of different wheat varieties under conventional fertilization treatment and the baseline yield under no fertilization treatment at different soil fertility levels in various regions, combined with the total amount of nitrogen applied.

[0013] The gene feature extraction module is used to perform genome sequencing on wheat varieties participating in the screening, identify and extract specific gene fragment features related to nutrient absorption, stress resistance and yield composition, and construct gene feature vectors.

[0014] The machine learning optimization control system integrates a data preprocessing engine, a feature fusion algorithm, and a deep learning classification model. It receives soil fertility level information, yield data, nitrogen fertilizer agronomic efficiency, and gene feature vectors to perform multi-dimensional data mining and nonlinear correlation modeling.

[0015] The decision output terminal is used to classify wheat varieties into high-yield and high-efficiency varieties, high-yield and low-efficiency varieties, low-yield and high-efficiency varieties, and low-yield and low-efficiency varieties under different soil fertility levels based on the calculation results of the machine learning optimization control system, and output corresponding optimization suggestions.

[0016] Preferably, when processing data, the basic productivity assessment server normalizes the yield data of different years and different ecological sites to eliminate random errors caused by climate fluctuations and obtain characteristic values ​​that reflect the essential productivity of arable land.

[0017] Furthermore, the regional soil fertility grading device executes differentiated judgment logic based on geographical divisions. In the northern region, the regional soil fertility grading device sets a first northern preset yield threshold and a second northern preset yield threshold. When the basic productivity is greater than or equal to the first northern preset yield threshold, it is judged as a high-fertility wheat field; when the basic productivity is between the second northern preset yield threshold and the first northern preset yield threshold, it is judged as a medium-fertility wheat field; and when the basic productivity is less than the second northern preset yield threshold, it is judged as a low-fertility wheat field.

[0018] Furthermore, in North China, the regional soil fertility classification device is set with a first preset yield threshold and a second preset yield threshold for North China; when the basic productivity is greater than or equal to the first preset yield threshold for North China, it is determined to be a high-fertility wheat field; when the basic productivity is between the second preset yield threshold for North China and the first preset yield threshold for North China, it is determined to be a medium-fertility wheat field; when the basic productivity is less than the second preset yield threshold for North China, it is determined to be a low-fertility wheat field.

[0019] Furthermore, in the middle and lower reaches of the Yangtze River, the regional soil fertility classification device is set with a first Yangtze River preset yield threshold and a second Yangtze River preset yield threshold; when the basic productivity is greater than or equal to the first Yangtze River preset yield threshold, it is determined to be a high-fertility wheat field; when the basic productivity is between the second Yangtze River preset yield threshold and the first Yangtze River preset yield threshold, it is determined to be a medium-fertility wheat field; when the basic productivity is less than the second Yangtze River preset yield threshold, it is determined to be a low-fertility wheat field.

[0020] Furthermore, in the southwest region, the regional soil fertility classification device is set with a first preset yield threshold and a second preset yield threshold for the southwest. When the basic productivity is greater than or equal to the first preset yield threshold for the southwest, it is determined to be a high-fertility wheat field. When the basic productivity is between the second preset yield threshold for the southwest and the first preset yield threshold for the southwest, it is determined to be a medium-fertility wheat field. When the basic productivity is less than the second preset yield threshold for the southwest, it is determined to be a low-fertility wheat field.

[0021] Furthermore, the nitrogen fertilizer utilization efficiency analysis platform uses a calculation logic described in pure Chinese text when calculating nitrogen fertilizer agronomic efficiency: that is, by obtaining the wheat yield per unit area under conventional fertilization treatment, subtracting the wheat yield per unit area under no fertilization treatment, the increased yield due to nitrogen fertilizer application is obtained. Then, the increased yield is divided by the mass of pure nitrogen input per unit area, and finally the yield gain that can be obtained by a unit nitrogen fertilizer input is obtained, that is, the nitrogen fertilizer agronomic efficiency value.

[0022] Furthermore, the data preprocessing engine in the machine learning optimization control system discretizes the gene sequence data provided by the gene feature extraction module, transforming the complex base arrangement into a numerical matrix suitable for input to the machine learning model.

[0023] Furthermore, the machine learning optimization control system utilizes algorithms such as support vector machines or gradient boosting decision trees to construct a variety selection model. This model is trained with soil fertility level and gene feature vectors as input variables, and expected yield and expected nitrogen fertilizer agronomic efficiency as target variables, to learn the deep mapping relationship between soil fertility level and wheat gene expression and nutrient utilization efficiency.

[0024] Preferably, when the machine learning optimization control system performs variety screening, it first calculates the average yield and average nitrogen fertilizer agronomic efficiency of all tested varieties under a specific soil fertility level in a specific region as a benchmark reference line.

[0025] Furthermore, the decision output terminal performs the following logical judgment: if the yield and nitrogen fertilizer agronomic efficiency of a certain wheat variety are both greater than the corresponding average, then the variety is marked as a high-yield and high-efficiency variety.

[0026] Furthermore, if a wheat variety has a yield greater than the average, but its nitrogen fertilizer agronomic efficiency is lower than the average, then the variety is marked as a high-yield, low-efficiency variety.

[0027] Furthermore, if the yield of a certain wheat variety is lower than the average, but the agronomic efficiency of nitrogen fertilizer is higher than the average, then the variety is marked as a low-yield, high-efficiency variety.

[0028] Furthermore, if the yield and nitrogen fertilizer agronomic efficiency of a certain wheat variety are both lower than the corresponding average, then the variety is marked as a low-yield and low-efficiency variety.

[0029] Furthermore, the gene feature extraction module uses high-throughput sequencing to focus on detecting gene families related to nitrogen transporters and amino acid metabolic pathways, and uses the detected gene polymorphism information as feature input to a machine learning model to predict the adaptability potential of varieties under different fertility environments.

[0030] Preferably, the soil fertility basic data acquisition terminal also includes a soil sensor cluster for collecting soil total nitrogen content, available phosphorus content, available potassium content and organic matter content. These physicochemical data are input as auxiliary factors into the basic productivity assessment server to correct the basic productivity values.

[0031] Furthermore, the machine learning optimization master control system employs cross-validation during training to continuously iterate and optimize the weight parameters in the model, thereby improving the system's prediction accuracy for unknown environments and newly bred varieties.

[0032] Furthermore, the decision output terminal not only outputs the variety classification results, but also provides the optimal nitrogen fertilizer application recommendation based on the specific soil fertility classification of the target wheat field through a prediction model.

[0033] Furthermore, the machine learning optimization control system is equipped with a knowledge base update module, which can periodically crawl the latest climate model data and agricultural technical indicators from external agricultural databases to realize the dynamic evolution of the prediction model.

[0034] Furthermore, the soil fertility basic data acquisition terminal has a data cleaning function, which can automatically identify and remove abnormal yield values ​​caused by severe pests and diseases or extreme weather, ensuring that the data input into the model is representative and scientific.

[0035] Preferably, when determining the soil fertility level, the regional soil fertility grading device will also refer to the average yield fluctuation rate of the plot over the previous three years. If the fluctuation rate exceeds the preset stability threshold, a multi-source data composite verification mechanism will be automatically triggered to ensure the accuracy of soil fertility grading.

[0036] Furthermore, the variety phenotypic monitoring equipment uses UAV remote sensing technology to collect vegetation indices at different growth stages of wheat, which are used as auxiliary indicators for yield prediction to compensate for the lack of real-time performance caused by relying solely on the final measured yield.

[0037] Furthermore, the machine learning optimization control system uses clustering analysis algorithms to classify wheat varieties with similar soil fertility response characteristics and discover common patterns in nutrient utilization among different varieties.

[0038] Furthermore, the optimal suggestions provided by the decision output terminal include the best spatiotemporal layout scheme for planting high-yield and high-efficiency varieties, guiding farmers to carry out precise sowing on plots with different soil fertility.

[0039] Compared with the prior art, the present invention has the following beneficial effects:

[0040] 1. This invention achieves precise coupling between wheat variety selection and soil fertility by constructing a machine learning-based optimization system. The system can automatically classify wheat fields into three soil fertility levels—high, medium, and low—based on specific thresholds for different ecological zones, overcoming the problems of simplistic and ambiguous soil fertility evaluation in traditional selection methods, thus making variety selection more regionally targeted and environmentally adaptable.

[0041] 2. This invention introduces a gene feature extraction module, combining macroscopic yield performance with microscopic gene characteristics, and utilizes machine learning models to uncover the genetic potential of varieties in nutrient absorption and transformation. This cross-scale information fusion improves the accuracy of the system's prediction of wheat nitrogen fertilizer use efficiency, providing a deep genetic basis for screening high-yield and high-efficiency varieties.

[0042] 3. This invention establishes a scientific nitrogen fertilizer agronomic efficiency evaluation system. Through textual logical calculations of yield, it accurately depicts the relationship between fertilizer input and output. The high-yield and high-efficiency varieties selected by the system can not only ensure grain yield but also reduce nitrogen fertilizer input costs, which is of great practical significance for achieving fertilizer reduction and efficiency improvement and the green transformation of agriculture.

[0043] 4. The system architecture of this invention boasts a high level of automation and intelligence. From basic data collection to hierarchical evaluation, and then to machine learning modeling and decision output, the entire process requires no manual intervention, greatly improving the efficiency of variety breeding and recommendation. Simultaneously, through dynamically updated knowledge bases and correction mechanisms, the system ensures its adaptability to constantly changing climatic conditions and agricultural production environments.

[0044] 5. This invention provides a refined variety classification strategy, detailing varieties into four dimensions, which can provide classification guidance for operators with different needs. For example, high-yield and high-efficiency varieties are recommended for wheat fields with high soil fertility to pursue maximum yield, while varieties with strong stress resistance and high nutrient utilization rate are recommended for wheat fields with low soil fertility, thereby achieving optimal allocation of agricultural production resources.

[0045] 6. This invention eliminates the rigid dependence on specific physical values ​​and empirical formulas, instead employing logically preset thresholds and textually described computational protocols, enhancing the system's universality and scalability. Whether in northern drylands or southern paddy fields, the system can quickly adapt to local wheat variety selection needs by adjusting internal logical parameters, demonstrating high technological transfer value. Attached Figure Description

[0046] Figure 1 This is a schematic diagram of the overall technical solution architecture according to the present invention;

[0047] Figure 2 This is a schematic diagram illustrating the core principle framework of multi-dimensional data mining and modeling of soil fertility, nitrogen fertilizer, and genes in the present invention.

[0048] Figure 3 This is a logical flowchart of the regional differentiated soil fertility classification and basic productivity assessment according to the present invention.

[0049] Figure 4 This is a schematic diagram of the multi-level interaction relationship and data flow between wheat gene feature extraction and nitrogen fertilizer use efficiency analysis according to the present invention.

[0050] Figure 5 This is a flowchart of the variety selection decision-making process based on both yield and efficiency criteria according to the present invention.

[0051] Figure 6 This is a logical framework diagram of the dynamic evolution of the knowledge base and the iterative optimization of the model according to the present invention. Detailed Implementation

[0052] Example 1: Please refer to the appendix Figure 1 To be continued Figure 6 To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to specific embodiments.

[0053] A machine learning-based wheat variety soil fertility nitrogen fertilizer gene detection and selection system includes a soil fertility basic data acquisition terminal, a basic productivity assessment server, a regional soil fertility grading device, a variety phenotypic monitoring device, a nitrogen fertilizer use efficiency analysis platform, a gene feature extraction module, a machine learning selection master control system, and a decision output terminal. The soil fertility basic data acquisition terminal establishes a bidirectional data communication connection with the basic productivity assessment server via an industrial-grade fieldbus or wireless sensor network. The basic productivity assessment server interacts with the regional soil fertility grading device via a high-speed data backplane. The variety phenotypic monitoring device, nitrogen fertilizer use efficiency analysis platform, and gene feature extraction module are all connected to the machine learning selection master control system via a data fusion gateway. The output of the machine learning selection master control system is connected to the decision output terminal for real-time distribution of variety selection instructions and classification results.

[0054] The soil fertility basic data acquisition terminal is used to acquire raw yield data of wheat under no-fertilization conditions, soil physicochemical property parameters, and historical tillage records within the target area, and transmit the above data to the basic productivity assessment server. At the hardware implementation level, the soil fertility basic data acquisition terminal includes a soil sensor cluster deployed in the sample plots to be tested. This cluster is embedded with high-precision nutrient sensors, moisture sensors, and conductivity sensors, and can collect data on total nitrogen content, available phosphorus content, available potassium content, and organic matter content in the soil in real time.

[0055] The soil fertility basic data acquisition terminal also integrates a handheld data entry terminal for receiving manually entered historical yield data and agricultural operation records. The terminal features data cleaning capabilities and a built-in outlier detection algorithm, automatically identifying and removing abnormal yield values ​​caused by sudden pest or disease outbreaks or extreme weather disasters, ensuring the data input for subsequent evaluation is highly representative and scientifically sound. During data preprocessing, the terminal standardizes the units of the raw yield data, converting yield values ​​from different units of measurement into tons per hectare.

[0056] The basic productivity assessment server is used to standardize the received wheat yield data under the no-fertilization treatment and calculate the basic productivity value of the cultivated land. The basic productivity value is defined as the wheat yield contributed by the fertility of the cultivated land soil itself, excluding interference from artificial fertilizer input. When processing the data, the basic productivity assessment server is configured to execute spatiotemporal normalization logic, mapping yield data collected from different years and different ecological sites to the same baseline dimension, eliminating random errors in yield caused by climate fluctuations, and obtaining steady-state characteristic values ​​reflecting the inherent productivity of the cultivated land. Furthermore, the basic productivity assessment server incorporates soil physicochemical data provided by a soil sensor cluster as an auxiliary factor, and corrects the initially calculated basic productivity value through a multiple linear regression model to compensate for the random bias of the no-fertilization yield data from a single year.

[0057] The regional soil fertility classification device uses pre-stored national farmland quality bulletin data, combined with the proportional distribution of high-quality, medium-quality, and low-quality farmland within a specific region, to compare the obtained basic productivity values ​​with preset soil fertility classification thresholds, thus classifying wheat fields in each region into high-fertility, medium-fertility, and low-fertility levels. The device integrates a geographic information system storage unit and pre-stores a dynamic threshold database for different administrative divisions and ecological zones.

[0058] The regional soil fertility grading device executes differentiated judgment logic based on geographical divisions: In the northern region, the device sets a first preset yield threshold and a second preset yield threshold; the first preset yield threshold is set to 4.5 tons per hectare, and the second preset yield threshold is set to 3.2 tons per hectare. When the calculated basic productivity value is greater than or equal to 4.5 tons per hectare, the device outputs a logic judgment signal for high-fertility wheat fields; when the basic productivity value is between 3.2 tons per hectare and 4.5 tons per hectare (inclusive of 3.2 tons per hectare but exclusive of 4.5 tons per hectare), it is judged as a medium-fertility wheat field; when the basic productivity value is less than 3.2 tons per hectare, it is judged as a low-fertility wheat field.

[0059] In North China, the regional soil fertility grading device is set with a first preset yield threshold and a second preset yield threshold for North China; wherein the first preset yield threshold for North China is set at 4.8 tons per hectare, and the second preset yield threshold for North China is set at 3.3 tons per hectare. When the basic productivity value is greater than or equal to 4.8 tons per hectare, it is determined to be a high-fertility wheat field; when the basic productivity value is between 3.3 tons per hectare and 4.8 tons per hectare, it is determined to be a medium-fertility wheat field; when the basic productivity value is less than 3.3 tons per hectare, it is determined to be a low-fertility wheat field.

[0060] In the middle and lower reaches of the Yangtze River, the regional soil fertility grading device is set with a first preset yield threshold and a second preset yield threshold for the Yangtze River; wherein, the first preset yield threshold is set to 3.4 tons per hectare, and the second preset yield threshold is set to 2.1 tons per hectare. When the basic productivity value is greater than or equal to 3.4 tons per hectare, it is determined to be a high-fertility wheat field; when the basic productivity value is between 2.1 tons per hectare and 3.4 tons per hectare, it is determined to be a medium-fertility wheat field; when the basic productivity value is less than 2.1 tons per hectare, it is determined to be a low-fertility wheat field.

[0061] In the southwest region, the regional soil fertility grading device is set with a first preset yield threshold and a second preset yield threshold; wherein, the first preset yield threshold is set to 3.1 tons per hectare, and the second preset yield threshold is set to 1.7 tons per hectare. When the basic productivity value is greater than or equal to 3.1 tons per hectare, it is determined to be a high-fertility wheat field; when the basic productivity value is between 1.7 tons per hectare and 3.1 tons per hectare, it is determined to be a medium-fertility wheat field; when the basic productivity value is less than 1.7 tons per hectare, it is determined to be a low-fertility wheat field.

[0062] When determining soil fertility levels, the regional soil fertility grading device also references the average yield fluctuation rate of the plot over the previous three years through its internal stability assessment unit. If the fluctuation rate exceeds a preset stability threshold, it indicates that the basic productivity of the plot is excessively affected by random environmental factors. In this case, the system will automatically trigger a multi-source data composite verification mechanism, retrieving historical yield records over a longer period or increasing the weight of soil physicochemical parameters for comprehensive secondary grading to ensure the accuracy of soil fertility classification.

[0063] The variety phenotypic monitoring equipment is used to monitor and record yield data of different wheat varieties under conventional fertilization treatments in real time on plots with different soil fertility levels. This equipment includes not only a traditional electronic weighing system but also a drone remote sensing inspection system. The drone remote sensing inspection system is equipped with a multispectral camera, which can collect electromagnetic wave reflection information in different bands during key growth stages such as wheat flowering and grain filling, calculate phenotypic indicators such as the Normalized Difference Vegetation Index (NDVI) and Enhanced Vegetation Index (EDI), and use these as auxiliary predictive factors for yield. This multi-dimensional monitoring method can capture the growth dynamics of varieties under different fertility environments in real time, overcoming the shortcomings of relying solely on yield measurements at harvest time in terms of data real-time performance and process description.

[0064] The nitrogen fertilizer use efficiency analysis platform is used to calculate the nitrogen fertilizer agronomic efficiency based on the yields of different wheat varieties under conventional fertilization treatments and the baseline yields under no-fertilization treatments at different soil fertility levels in various regions, combined with the total amount of nitrogen applied. The platform internally stores a nitrogen fertilizer efficiency calculation protocol, which uses a calculation logic described in pure Chinese text: First, it obtains the yield data per unit area of ​​a specific wheat variety under conventional fertilization treatments and defines it as the fertilization yield value; then, it obtains the baseline yield data per unit area of ​​the corresponding plot under no-fertilization treatments and defines it as the control yield value; next, it calculates the difference between the fertilization yield value and the control yield value, which represents the net yield increase due to nitrogen fertilizer application; finally, it divides this net yield increase by the actual mass of pure nitrogen applied per unit area, and the resulting quotient is the nitrogen fertilizer agronomic efficiency value. This value directly reflects the wheat yield gain obtained per unit mass of nitrogen fertilizer applied and is a key indicator for evaluating the efficiency of variety resource utilization.

[0065] Specifically, the calculation protocol of the nitrogen fertilizer utilization efficiency analysis platform is rigorously implemented mathematically. The formula for calculating the nitrogen fertilizer agronomic efficiency is:

[0066]

[0067] in, This indicates the yield of wheat per unit area under conventional fertilization treatment (unit: tons / hectare). This represents the basic yield per unit area of ​​the corresponding plot under no-fertilization treatment (unit: tons / hectare), a value provided by the basic productivity assessment server. This represents the actual mass of pure nitrogen input per unit area (unit: kg / ha). Calculation results The unit is kilograms of grains / kilogram of pure nitrogen.

[0068] The dynamic physiological utilization rate calculation module calculates the apparent nitrogen utilization rate ( The formula for ) is:

[0069]

[0070] in, and The figures represent the total nitrogen accumulation of mature plants under fertilized and unfertilized conditions (unit: kg / ha). This data was measured periodically during the wheat grain-filling stage using a portable chlorophyll meter or near-ground spectrometer mounted on a field phenotyping platform, and analyzed using a built-in empirical model. The estimate was obtained. The formula shows that... It reflects the efficiency of the plant in absorbing nitrogen from the applied fertilizer.

[0071] The gene feature extraction module is used to perform genome sequencing on wheat varieties participating in the screening, identifying and extracting specific gene fragment features related to nutrient uptake, stress resistance, and yield composition. This module integrates a high-throughput sequencing unit, enabling whole-genome resequencing or target region capture for the complex wheat genome. The module focuses on detecting gene families related to nitrogen transporter families, amino acid metabolic pathways, and root development regulation, identifying single nucleotide polymorphism sites or insertion / deletion markers. Subsequently, the module converts the detected genotype data into digital expressions, constructing gene feature vectors. These vectors can characterize the sensitivity and transformation potential of wheat varieties to fertilizer at the micro-genetic level, providing deep genetic input for subsequent machine learning modeling.

[0072] In a preferred embodiment, the gene feature extraction module digitally encodes single nucleotide polymorphism (SNP) sites obtained from high-throughput sequencing. For each SNP site, a counting encoding method is used based on its genotype (e.g., AA, AT, TT): homozygous major allele is encoded as 0, heterozygous genotype as 1, and homozygous minor allele as 2. If there are a total of If there are 1 SNP locus, then the variety can be represented as: 3D integer vector .

[0073] To reduce dimensionality and remove redundancy, the following is applied: The original SNP matrix composed of varieties (dimensions) Perform principal component analysis (PCA) for dimensionality reduction. The specific steps are as follows:

[0074] Centering the original matrix yields the matrix .

[0075] Calculate the covariance matrix: .

[0076] For covariance matrix Perform eigenvalue decomposition to obtain eigenvalues. and the corresponding feature vectors .

[0077] Before choosing The eigenvectors corresponding to the eigenvalues ​​constitute the projection matrix. , The choice makes the former The cumulative variance contribution of each principal component exceeds 85%.

[0078] Ultimately, the gene feature vector for each variety Through formula The dimension of the result is This vector serves as part of the input features for the machine learning optimization master system.

[0079] The machine learning optimization control system, serving as the intelligent core of the entire system, integrates a data preprocessing engine, a feature fusion algorithm, and a deep learning classification model. The data preprocessing engine discretizes and reduces the dimensionality of the raw gene sequence data provided by the gene feature extraction module, transforming massive base permutations into a high-dimensional numerical matrix suitable for deep learning model input. The feature fusion algorithm is configured to perform feature space alignment and weighted fusion of heterogeneous data such as soil fertility level, soil physicochemical parameters, vegetation index, gene feature vectors, and conventional fertilization yield, constructing a multi-dimensional input feature set.

[0080] The machine learning optimization control system utilizes algorithms such as support vector machines, gradient boosting decision trees, or convolutional neural networks to construct variety selection models. During the training phase, the model uses soil fertility level and gene feature vectors as the main input variables, and expected yield and expected nitrogen fertilizer agronomic efficiency as target variables. Through cross-validation, the model iteratively learns on training sets in multiple ecological zones, continuously optimizing internal weight parameters and bias terms, establishing a non-linear mapping relationship between soil fertility level and wheat gene expression and nutrient use efficiency. Furthermore, the machine learning optimization control system is equipped with a knowledge base update module, which can periodically crawl the latest climate change data, pest and disease trends, and new breeding technology indicators from external agricultural research databases via a network interface, enabling dynamic evolution and self-reinforcement of the prediction model.

[0081] In one specific embodiment, the machine learning optimization control system uses an integrated model of convolutional neural network (CNN) and gradient boosting decision tree (GBDT) for variety selection, and its specific implementation is as follows:

[0082] Input feature tensor: , dimension .in, The sample size is (number of combinations of variety × soil fertility grade). The fused feature dimensions consist of soil fertility level (unique thermal encoding, 3-dimensional), soil physicochemical parameters (normalized, 4-dimensional), vegetation index (normalized, 2-dimensional), and gene feature vector (PCA dimensionality reduction). It is a composite of (1-dimensional) and conventional fertilization yield (normalized, 1-dimensional), therefore .

[0083] Target variable vector: , dimension It includes two regression objectives: expected output. and expected nitrogen fertilizer agronomic efficiency .

[0084] The system first uses a 1D convolutional neural network to process the input feature tensor. This is to capture local correlations between features. The CNN module contains two convolutional blocks. Each convolutional block consists of a one-dimensional convolutional layer, a batch normalization layer, and a ReLU activation function. The first convolutional block has 32 kernels and a kernel size of 3; the second convolutional block has 64 kernels and a kernel size of 3. The convolutional layers are followed by a global max pooling layer, compressing the feature map into a 64-dimensional feature vector. .

[0085] Subsequently, the feature vectors extracted by the CNN Compared with the original input features The features are concatenated to form an enhanced feature vector. The vector is fed into a regressor consisting of Gradient Boosting Decision Trees (GBDT). GBDT uses squared error as the splitting criterion, iteratively trains 100 decision trees, each with a maximum depth of 5 and a learning rate of 0.1.

[0086] The model training employs a multi-task learning strategy, simultaneously predicting yield and nitrogen fertilizer agronomic efficiency. The total loss function is a weighted sum of the losses from the two tasks:

[0087]

[0088] in, and These represent the mean squared error losses for the yield prediction task and the nitrogen fertilizer efficiency prediction task, respectively. and To balance the importance of the two tasks, all hyperparameters are set to 1.0 in this embodiment. Mean squared error loss. The calculation formula is:

[0089]

[0090] in, For the first The true label value of each sample These are the model's predicted values. The model uses the Adam optimizer to update parameters, with an initial learning rate of 0.001, and employs a cosine annealing strategy to dynamically adjust the learning rate.

[0091] The decision output terminal is used to classify wheat varieties in a refined manner and output corresponding optimization suggestions under different soil fertility levels based on the calculation results of the machine learning optimization control system. During logical judgment, the machine learning optimization control system first calculates the average yield and average nitrogen fertilizer agronomic efficiency of all tested wheat varieties under a specific soil fertility level in a specific region, and sets these two averages as the benchmark reference line for variety selection.

[0092] The decision output terminal executes the following specific classification logic judgments: A. If the yield and nitrogen fertilizer agronomic efficiency values ​​of a certain wheat variety are both greater than the corresponding system average, then the variety is marked as a high-yield and high-efficiency variety, and a first-type recommendation strategy is generated, suggesting large-scale promotion in wheat fields with high and medium fertility to pursue maximum yield and optimal fertilizer utilization; B. If the yield of a certain wheat variety is greater than the average, but the nitrogen fertilizer agronomic efficiency value is lower than the average, then the variety is marked as a high-yield and low-efficiency variety, and a second-type recommendation strategy is generated, suggesting planting in wheat fields with excellent fertility and appropriately increasing fertilizer input to maintain its high-yield characteristics; C. If the yield of a certain wheat variety is lower than the average, but the nitrogen fertilizer agronomic efficiency value is higher than the average, then the variety is marked as a low-yield and high-efficiency variety, and a third-type recommendation strategy is generated, suggesting promotion in wheat fields with low fertility or arid and barren areas to utilize its excellent resource acquisition capabilities to achieve stable yield and increased efficiency; D. If the yield and nitrogen fertilizer agronomic efficiency values ​​of a certain wheat variety are both lower than the corresponding average values, the variety will be marked as a low-yield and low-efficiency variety, and a recommendation to eliminate it will be issued or it will be recommended to use it as a negative control in a specific breeding study.

[0093] The decision output terminal executes the following classification logic, the core of which is to calculate the relative performance index of each variety and compare it with a threshold:

[0094] First, calculate the average yield of all tested varieties under a specific region and specific soil fertility level. and average nitrogen fertilizer agronomic efficiency :

[0095]

[0096] in, This represents the total number of varieties within this group. and The first Measured yield and nitrogen fertilizer agronomic efficiency of each variety.

[0097] Then, a comprehensive excellence score is calculated for each variety. :

[0098]

[0099] in, and The yield and efficiency values ​​for the variety to be determined; and These are the standard deviations of the output and efficiency values ​​within the group, respectively. and These are weighting coefficients, with a default value of 1.0, which can be adjusted according to policy objectives.

[0100] The specific rules for classifying varieties are as follows:

[0101] like and If it is, then it is determined to be a high-yield and high-efficiency variety.

[0102] like and If it is, it is judged to be a high-yield but low-efficiency variety.

[0103] like and If it is, it is determined to be a low-yield, high-efficiency variety.

[0104] like and If so, it is judged as a low-yield and low-efficiency variety.

[0105] For high-yield and high-efficiency varieties, the system will further output planting suggestions. The optimal spatiotemporal layout scheme for planting is generated by an independent submodule. This submodule calculates the normalized score of the variety for each local fertility level and each sowing window based on the variety's historical performance data, using the min-max normalization method, and recommends the combination scheme with the highest score.

[0106] The optimal suggestions provided by the decision output terminal further include the best spatiotemporal layout scheme for high-yield and high-efficiency varieties. This scheme combines climate model prediction models to provide suggested sowing windows, appropriate basic seedling numbers, and recommended dosages of nitrogen fertilizer applied in stages. By guiding farmers to carry out precision sowing and integrated water and fertilizer management, it maximizes the genetic yield potential of the varieties.

[0107] Example 2: Based on Example 1, this example describes a system implementation method based on edge computing and distributed architecture, which aims to improve the real-time performance of data processing and the system disaster recovery capability in cross-provincial large-scale variety screening applications.

[0108] In this embodiment, the soil fertility basic data acquisition terminal is configured as an edge computing node with local preprocessing capabilities. Each sample plot is equipped with an edge gateway, which integrates a low-power, high-performance microprocessor capable of performing instantaneous filtering of the soil physicochemical data collected by sensors locally. The edge gateway employs a moving average filtering algorithm to eliminate high-frequency noise during sensor sampling. When the soil fertility basic data acquisition terminal detects a sudden and drastic fluctuation in basic productivity, the edge gateway can autonomously trigger a high-frequency sampling mode and send an early warning signal to the basic productivity assessment server.

[0109] The basic productivity assessment server employs a distributed server cluster architecture, using a load balancer to distribute massive data collection tasks across different physical nodes. When calculating basic productivity values, the server cluster incorporates a textual implementation of the spatial kriging interpolation algorithm. This involves analyzing the spatial autocorrelation between known sampling points and using a semi-variogram to estimate the productivity characteristics of unknown areas. This approach addresses assessment biases caused by uneven distribution of sampling points across vast arable land.

[0110] In this embodiment, the regional soil fertility grading device is designed as a dynamic grading unit based on a rule engine. This unit not only stores preset yield thresholds for four major regions—North China, the middle and lower reaches of the Yangtze River, and Southwest China—but also automatically adjusts these thresholds based on real-time meteorological monitoring data. For example, in years of extreme drought, the regional soil fertility grading device automatically triggers a threshold reduction protocol, lowering the criteria for identifying high-fertility wheat fields by a preset ratio to ensure that the soil fertility grading results conform to the actual ecological carrying capacity of the season.

[0111] In Example 2, the variety phenotypic monitoring equipment was enhanced with an automated field phenotypic platform. This platform includes a multi-sensor imaging system deployed on a track, capable of monitoring wheat plant height, ear number, and leaf area index around the clock. These microscopic phenotypic parameters are transmitted in real-time to a nitrogen fertilizer use efficiency analysis platform via high-speed industrial Ethernet. In this example, the nitrogen fertilizer use efficiency analysis platform incorporates a dynamic physiological utilization rate calculation module, which not only calculates the final agronomic efficiency but also assesses the variety's nitrogen translocation efficiency at different growth stages by analyzing the dynamic balance of nitrogen concentration within the plant. The calculation logic is described as follows: obtain the total nitrogen accumulation of mature plants, subtract the total nitrogen accumulation of plants under no-fertilizer treatment, and then divide by the total nitrogen applied to obtain the apparent nitrogen utilization rate.

[0112] In Example 2, the gene feature extraction module employs a cloud computing collaborative mode. The massive amounts of raw data (FASTQ format) generated by sequencing are uploaded to a gene feature cloud server via a dedicated line, where parallel alignment and variant detection algorithms are executed. This server can identify specific haplotype blocks associated with nitrogen efficiency and map this polymorphic information into unified digital tags, sending them back to the machine learning optimization master control system.

[0113] The machine learning optimization control system incorporates an ensemble learning strategy at the algorithm level. This system simultaneously runs multiple heterogeneous sub-models, including a random forest model, an extreme gradient boosting model, and a deep multilayer perceptron model. Each sub-model independently predicts the yield and efficiency of a variety. The machine learning optimization control system includes a voting mechanism or weighted averaging unit to non-linearly fuse the outputs of multiple sub-models, eliminating inductive bias that might arise from a single algorithm. During training, the system utilizes L2 regularization and Dropout techniques to prevent overfitting on small gene datasets.

[0114] In this embodiment, the decision output terminal is expanded into a multi-terminal synchronous interactive system. Besides providing detailed variety evaluation reports to breeding experts, it also connects to farmers' mobile applications via the mobile internet. Based on the geographical coordinates input by the farmer, the terminal can automatically match the soil fertility level and, based on machine learning model predictions, provide a list of the most suitable advantageous varieties for that plot of land. Simultaneously, it outputs a personalized fertilization curve based on the nitrogen fertilizer response characteristics of that variety.

[0115] Furthermore, the system in Example 2 also includes an environmental stress monitoring module. This module monitors the temperature, humidity, light radiation, and pest and disease indices of the sample plots in real time, and inputs these environmental covariates into the machine learning optimization master control system. By introducing covariance structure analysis, the system can isolate the contribution of environmental stress to yield formation, more accurately extract the genetic stability parameters of the variety itself, and ensure that the selected high-yield and high-efficiency varieties have stable performance in different years.

[0116] Example 3: This example describes a system architecture that integrates an automated laboratory with high-throughput genotyping technology. It is mainly applied to national-level wheat breeding bases and emphasizes closed-loop optimization from microscopic molecular breeding to macroscopic soil fertility adaptation.

[0117] In this embodiment, the gene feature extraction module is specifically implemented as a high-throughput molecular marker genotyping pipeline. This pipeline includes an automated DNA extraction workstation, a pipetting robot, and a gene chip scanner. Using a customized wheat 66K or higher-density SNP chip, this module can rapidly genotype tens of thousands of varietal samples. The extracted features include not only nutrient uptake-related genes but also alleles related to rust resistance, Fusarium head blight resistance, and stress tolerance (such as heat and cold tolerance). All this gene information is converted into a high-dimensional sparse matrix, which is then dimensionality-reduced using principal component analysis and input into a machine learning optimization control system.

[0118] The soil fertility basic data acquisition terminal in Example 3 integrates a hyperspectral geoscience analysis unit. This unit can directly retrieve core fertility parameters such as soil organic matter, total nitrogen, and cation exchange capacity through hyperspectral scanning of exposed topsoil. Its internal retrieval algorithm establishes a textual correspondence between spectral reflectance and physical and chemical values ​​through first-order differential processing and band combination. This allows soil fertility assessment to go beyond historical yields and gain in-depth interpretability at the chemical level.

[0119] The basic productivity assessment server is configured as a digital twin system based on crop growth simulation models (such as DSSAT or APSIM models). This server constructs a virtual wheat growing environment using collected soil physicochemical parameters and historical meteorological data. By simulating vegetation growth under unfertilized conditions in digital space, the system can generate a virtual control yield, which is then weighted and averaged with the measured yield. This digital twin technology enhances the robustness of basic productivity calculations under complex climatic conditions.

[0120] In this embodiment, the regional soil fertility grading device employs fuzzy comprehensive evaluation logic. Instead of rigidly classifying based solely on a single yield threshold, the system uses basic productivity, soil texture index, irrigation guarantee rate, and soil thickness as inputs, and utilizes fuzzy membership functions to classify wheat fields into high, medium, and low soil fertility probability distributions. This probabilistic grading method more accurately reflects the continuous changes in soil fertility levels, avoiding jumps in variety recommendation results at threshold boundaries.

[0121] The variety phenotypic monitoring equipment integrates an in-situ root monitoring device. This device, through a transparent observation tube buried beneath the sample plot and an automatic scanning imaging head, acquires real-time data on wheat root growth rate, branching density, and root depth distribution. Since nitrogen fertilizer absorption is closely related to root architecture, this underground phenotypic data is used to model a correlation between the deep neural network of the machine learning optimization control system and the aboveground yield data, identifying superior genotypes with strong root systems and efficient translocation characteristics.

[0122] In this embodiment, the machine learning optimization control system employs a transfer learning-based architecture. The system is pre-trained on a large historical database containing hundreds of thousands of wheat variety performance data, learning the general patterns of wheat growth and environmental interaction. Subsequently, it is fine-tuned using a small amount of measured data from a specific ecological zone. This strategy addresses the problem of low model prediction accuracy due to insufficient data sample size in the early stages of promoting newly bred varieties.

[0123] In this embodiment, the decision output terminal possesses multi-objective programming capabilities. While providing recommendations for preferred varieties, the terminal can search for optimal variety combinations and resource allocation schemes using genetic algorithms based on set agricultural policy objectives (such as maximizing grain yield or achieving zero growth in fertilizer use). For example, within a specific county, it can determine how to allocate the planting ratio of high-yield, high-efficiency varieties to low-yield, high-efficiency varieties to ensure total yield while maximizing the nitrogen fertilizer input-output ratio for the entire county.

[0124] The system is also equipped with a traceable quality tracking module. All collected soil fertility data, gene sequencing results, monitoring process images, and the final decision-making logic are recorded in a distributed ledger based on blockchain technology. This ensures the scientific rigor, impartiality, and traceability of each variety selection result, providing authoritative data support for seed companies' variety promotion and government agricultural technology subsidies.

[0125] Example 4: This example focuses on describing a variant of the system architecture for adapting to extreme environments and climate change, which particularly enhances the ability to analyze the coupling relationship between stress resistance genes and geodynamics.

[0126] In this embodiment, the soil fertility data acquisition terminal adds the acquisition of soil physical structure characteristics, such as soil bulk density, porosity, and infiltration rate. These physical indicators have been shown to significantly affect nitrogen mobility in the soil profile. The acquisition terminal converts these physical parameters into a structured data stream and transmits it to the basic productivity assessment server.

[0127] The basic productivity assessment server incorporates potential limiting factor analysis logic. When calculating basic productivity, the server automatically identifies the core bottlenecks leading to low yields. For example, if a plot of land has extremely low yields without fertilization, the system compares soil sensor data to determine whether it's due to nutrient deficiency, excessive heavy metals, or severe salinization. If it's caused by environmental stress, the plot's ISP value will be specifically adjusted, and it will be labeled as stress-related low soil fertility.

[0128] The regional soil fertility classification device correspondingly adds an adversity level dimension. In addition to classifying soil fertility as high, medium, and low, it also marks attributes such as drought-prone areas and saline-alkali improvement areas. For these special plots, the judgment threshold of the regional soil fertility classification device will be dynamically adjusted twice based on the local improvement history.

[0129] In this embodiment, the gene feature extraction module focuses on detecting transcription factor genes related to abiotic stress responses. For example, it extracts polymorphic features of genes related to abscisic acid signal transduction, accumulation of osmotic regulatory substances, and activity of antioxidant enzyme systems in the variety. This module combines these stress-response gene features with nitrogen fertilizer utilization gene features to generate a synergistic feature vector.

[0130] The machine learning optimization control system employs a multi-task learning network model. This model simultaneously predicts three tasks: Task 1 is the expected yield under different soil fertility levels; Task 2 is the nitrogen fertilizer agronomic efficiency; and Task 3 is the yield reduction rate (i.e., stability coefficient) of varieties under extreme high temperatures or drought. Through parameter sharing among tasks, the model can learn the characteristics of robust, superior varieties that can efficiently utilize nitrogen fertilizer and maintain stable yields under adverse conditions.

[0131] When outputting recommendations, the decision-making output terminal incorporates risk assessment weights. For risk-averse operators (such as smallholders), the system prioritizes varieties with high stability coefficients and low dependence on soil fertility; for operators with high input-output capabilities (such as agricultural cooperatives), varieties with extremely high yield ceilings are recommended. An environmental adaptability score is added to the decision-making logic: the expected yield reduction rate of a variety is multiplied by the probability of extreme weather events in the region to calculate a risk value. If the risk value exceeds a preset safety threshold, even if its yield and efficiency are both greater than the average, it will not be marked as a first-category recommended variety.

[0132] In this embodiment, the knowledge base update module of the machine learning optimization control system interfaces with a satellite remote sensing big data platform. This module crawls wheat performance data from similar ecological zones globally in real time, and uses cross-border similar environmental data to assist in correcting the local model, thus expanding the geographical adaptability boundary of the system.

[0133] In summary, the above embodiments, through multi-dimensional hardware integration and advanced machine learning algorithms, achieve precise coupling between wheat variety selection and soil fertility. The system, through in-depth mining of ISP, NAE, and genetic characteristics, transforms traditional, fuzzy, experience-based selection into precise, digital assessment. Whether in the North China Plain with its complex soil fertility distribution or in the ecologically diverse Yangtze River basin, the system can dynamically adjust logical thresholds and model parameters to output the most regionally targeted variety selection decisions. This end-to-end, intelligent technical solution enhances the scientific rigor of breeding and grain cultivation, and provides strong technical support for achieving reduced fertilizer use and increased efficiency in agriculture, as well as ensuring food security.

[0134] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention. Any obvious improvements, component substitutions, or adjustments to logical parameters without departing from the concept of the present invention, whose technical effects do not exceed the essential scope of the present invention, should be considered as extended embodiments of the present invention. Those skilled in the art can reasonably physically integrate or logically decompose the modules in the system according to the hardware performance, data scale, and application accuracy requirements of the actual production environment; these actions do not affect the integrity of the technical solution of the present invention or the validity of the claims.

Claims

1. A wheat variety soil fertility nitrogen fertilizer gene detection and optimization system based on machine learning, characterized in that, include: The soil fertility basic data acquisition terminal is used to acquire raw yield data of wheat under no-fertilization treatment conditions, soil physicochemical property parameters, and historical tillage records within the target area; The basic productivity assessment server is connected in communication with the soil fertility basic data acquisition terminal. It is used to standardize the received wheat yield data under the no-fertilization treatment and calculate the basic productivity value of the cultivated land. A regional soil fertility classification device, connected to the basic productivity assessment server, is used to compare the acquired basic productivity values ​​with preset soil fertility classification thresholds to classify wheat fields in each region into high soil fertility level, medium soil fertility level, and low soil fertility level. Variety phenotypic monitoring equipment is used to monitor and record yield data of different wheat varieties under conventional fertilization treatments in real time on sample plots with different soil fertility levels. The nitrogen fertilizer use efficiency analysis platform is used to calculate the nitrogen fertilizer agronomic efficiency based on the yield under conventional fertilization treatment, the baseline yield under no fertilization treatment, and the total amount of nitrogen applied. The gene feature extraction module is used to identify and extract gene fragment features related to nutrient absorption in wheat varieties and construct gene feature vectors. The machine learning optimization control system is connected to the regional soil fertility grading device, variety phenotyping equipment, nitrogen fertilizer use efficiency analysis platform and gene feature extraction module, respectively. It performs nonlinear correlation modeling by receiving soil fertility level, yield data, nitrogen fertilizer agronomic efficiency and gene feature vectors. The decision output terminal is connected to the machine learning optimization control system and is used to classify wheat varieties and output optimization suggestions based on the modeling calculation results.

2. The wheat variety soil fertility nitrogen fertilizer gene detection and optimization system based on machine learning according to claim 1, characterized in that: The soil fertility basic data acquisition terminal includes a soil sensor cluster deployed in the sample plot to be tested. The soil sensor cluster is embedded with a high-precision nutrient sensor, a moisture sensor and a conductivity sensor, which are used to collect the total nitrogen content, available phosphorus content, available potassium content and organic matter content of the soil in real time. The soil fertility basic data acquisition terminal has a data cleaning function, which automatically identifies and removes abnormal yield values ​​caused by pests and diseases or extreme weather through the built-in outlier detection algorithm, and performs unit unification processing on the original yield data in the data preprocessing stage. The ground fertility data acquisition terminal also includes an edge gateway with local preprocessing capabilities. The edge gateway integrates a microprocessor and is used to eliminate high-frequency noise during the sensor sampling process using a moving average filtering algorithm. When the monitored basic productivity experiences a sudden and drastic fluctuation, the edge gateway autonomously triggers a high-frequency sampling mode and sends an early warning signal to the basic productivity assessment server.

3. The wheat variety soil fertility nitrogen fertilizer gene detection and optimization system based on machine learning according to claim 2, characterized in that: The basic productivity assessment server adopts a distributed server cluster architecture, allocates collection tasks through a load balancer, and executes spatiotemporal normalization logic when calculating basic productivity values. This maps the yield data collected from different years and different ecological sites to the same benchmark dimension, eliminates the random error in yield formation caused by climate fluctuations, and obtains steady-state characteristic values ​​that reflect the essential productivity of arable land. The basic productivity assessment server also incorporates soil physicochemical data provided by the soil sensor cluster as an auxiliary factor, and corrects the preliminary calculated basic productivity values ​​through a multiple linear regression model to compensate for the random bias of data from a single year. The basic productivity assessment server is also configured as a digital twin system based on a crop growth simulation model. It uses soil physicochemical parameters and historical meteorological data to construct a virtual wheat growth environment, simulates the vegetation growth process under no-fertilization conditions to generate a virtual control yield, and performs a weighted average with the measured yield.

4. The wheat variety soil fertility nitrogen fertilizer gene detection and optimization system based on machine learning according to claim 3, characterized in that: The regional soil fertility classification device executes differentiated judgment logic based on geographical zoning. In the northern region, the regional soil fertility classification device is set with a first northern preset yield threshold and a second northern preset yield threshold, wherein the first northern preset yield threshold is greater than the second northern preset yield threshold. When the basic productivity is greater than or equal to the first preset yield threshold in the north, it is determined to be a high-fertility wheat field; When the basic productivity is between the second northern preset production threshold and the first northern preset production threshold, it is determined to be a wheat field with medium productivity. When the basic productivity is less than the second preset yield threshold in the north, it is judged as a low soil productivity wheat field; In North China, the middle and lower reaches of the Yangtze River, and Southwest China, the regional soil fertility classification device is respectively set with a first threshold and a second threshold, and makes a logical judgment of high soil fertility, medium soil fertility, and low soil fertility based on the relationship between the basic productivity value and the corresponding regional threshold. When determining the soil fertility level, the regional soil fertility grading device also refers to the average yield fluctuation rate of the corresponding plot in the previous three years through the internal stability assessment unit. If the fluctuation rate exceeds the preset stability threshold, the multi-source data composite verification mechanism is automatically triggered to retrieve historical yield records and add soil physicochemical parameter weights for comprehensive secondary grading.

5. The wheat variety soil fertility nitrogen fertilizer gene detection and optimization system based on machine learning according to claim 4, characterized in that: The variety phenotypic monitoring equipment includes an electronic weighing system and an unmanned aerial vehicle (UAV) remote sensing inspection system. The UAV remote sensing inspection system is equipped with a multispectral camera to collect electromagnetic wave reflection information of different bands at key stages of wheat growth, calculate the normalized vegetation index and the enhanced vegetation index, and input the enhanced vegetation index as an auxiliary predictor of yield into the machine learning optimization master control system. The variety phenotypic monitoring equipment also integrates an automated field phenotypic platform, which includes a multi-sensor imaging system deployed on a track for all-weather monitoring of wheat plant height, ear number, and leaf area index. The variety phenotypic monitoring equipment also includes an in-situ root monitoring device, which uses a transparent observation tube buried under the sample plot and an automatic scanning imaging head to acquire data on the growth rate, branching density, and root depth distribution of wheat roots in real time.

6. The wheat variety soil fertility nitrogen fertilizer gene detection and optimization system based on machine learning according to claim 5, characterized in that: The nitrogen fertilizer utilization efficiency analysis platform internally stores a nitrogen fertilizer efficiency calculation protocol. The calculation protocol uses a text-described operation logic: First, obtain the unit area fertilizer yield value of wheat varieties under conventional fertilization treatment. Subsequently, the unit area control yield value of the corresponding plot under the no-fertilizer treatment was obtained; Next, the difference between the fertilizer yield per unit area and the control yield is calculated, and this difference is defined as the net increase in yield resulting from the application of nitrogen fertilizer. Finally, the net increase in yield is divided by the actual amount of pure nitrogen input per unit area, and the resulting quotient is the nitrogen fertilizer agronomic efficiency value. The nitrogen fertilizer utilization efficiency analysis platform is also equipped with a dynamic physiological utilization rate calculation module, which is used to obtain the total nitrogen accumulation of mature plants, subtract the total nitrogen accumulation of plants under no-fertilization treatment, and then divide by the total amount of nitrogen applied to calculate the apparent nitrogen utilization rate, which is used to evaluate the nitrogen translocation efficiency of varieties at different growth stages.

7. The wheat variety soil fertility nitrogen fertilizer gene detection and optimization system based on machine learning according to claim 6, characterized in that: The gene feature extraction module integrates a high-throughput sequencing unit for resequencing the wheat genome, focusing on detecting gene families related to nitrogen transporter families, amino acid metabolic pathways, and root development regulation, and identifying single nucleotide polymorphism sites and insertion / deletion markers within them. The gene feature extraction module converts the detected genotype data into digital expressions and constructs gene feature vectors; In the case where the gene feature extraction module is specifically implemented as a high-throughput molecular marker typing pipeline, it includes an automated deoxyribonucleic acid extraction workstation, a pipetting robot, and a gene chip scanner, which rapidly types the variety samples by using a high-density single nucleotide polymorphism chip; The extracted features also include allele information related to resistance to rust, resistance to Fusarium head blight, and heat and cold stress tolerance. These features are then transformed into a high-dimensional sparse matrix, which is then dimensionality-reduced using principal component analysis and input into the machine learning optimization master control system.

8. The wheat variety soil fertility nitrogen fertilizer gene detection and optimization system based on machine learning according to claim 7, characterized in that: The machine learning optimization control system integrates a data preprocessing engine, a feature fusion algorithm, and a deep learning classification model. The data preprocessing engine discretizes and reduces the dimensionality of the original gene sequence data, transforming the base arrangement into a numerical matrix. The feature fusion algorithm performs feature space alignment and weighted fusion of soil fertility level, soil physicochemical parameters, vegetation index, gene feature vector, and conventional fertilizer yield. The machine learning optimization control system uses support vector machines, gradient boosting decision trees or convolutional neural network algorithms to build a variety selection model, and uses cross-validation to iteratively learn on training sets in multiple ecological zones to optimize weight parameters and bias terms. The machine learning optimization control system is also equipped with a knowledge base update module, which is used to regularly crawl the latest climate change data, pest and disease trends, and breeding technology indicators from external agricultural research databases. During the training process, the machine learning optimization master control system also uses regularization techniques and random neuron dropping techniques to prevent the model from overfitting on small sample gene datasets.

9. The wheat variety soil fertility nitrogen fertilizer gene detection and optimization system based on machine learning according to claim 8, characterized in that: The decision output terminal performs the following classification logic judgment: First, it calculates the average yield and average nitrogen fertilizer agronomic efficiency of all tested wheat varieties under the local soil fertility level as a baseline reference line; If the yield and nitrogen fertilizer agronomic efficiency values ​​of a certain wheat variety are both greater than the corresponding benchmark reference line, then the variety is marked as a high-yield and high-efficiency variety, and the first type of recommendation strategy is generated. If the yield value of a certain wheat variety is greater than the corresponding benchmark reference line, but the nitrogen fertilizer agronomic efficiency value is lower than the corresponding benchmark reference line, then the variety is marked as a high-yield and low-efficiency variety, and a second type of recommendation strategy is generated. If the yield of a certain wheat variety is lower than the corresponding benchmark reference line, but the nitrogen fertilizer agronomic efficiency value is higher than the corresponding benchmark reference line, then the variety is marked as a low-yield and high-efficiency variety, and a third type of recommendation strategy is generated. If the yield and nitrogen fertilizer agronomic efficiency values ​​of a certain wheat variety are both lower than the corresponding benchmark reference line, the variety will be marked as a low-yield and low-efficiency variety, and a recommendation to eliminate it will be output. The preferred recommendations provided by the decision output terminal also include optimal spatiotemporal layout schemes for planting high-yield and high-efficiency varieties.

10. The wheat variety soil fertility nitrogen fertilizer gene detection and optimization system based on machine learning according to claim 9, characterized in that: The system also includes an environmental stress monitoring module, which is used to monitor the temperature, humidity, light radiation and pest index of the sample plot in real time, and input the environmental covariates into the machine learning optimization master control system. Through covariance structure analysis, the contribution of environmental stress to yield formation is extracted, and the genetic stability parameters of the variety are extracted. The decision output terminal also has a multi-objective programming function, which is used to search for the optimal combination of varieties and resource allocation schemes through genetic algorithms according to the set agricultural policy objectives. In addition, the system is also equipped with a traceable quality traceability module, which records the collected soil fertility data, gene sequencing results, monitoring process images and the final decision-making logic in a distributed ledger based on blockchain technology to ensure the traceability of variety selection results. The machine learning optimization control system is also equipped with potential limiting factor analysis logic, which is used to compare soil sensor data when calculating basic productivity to determine whether the core weakness causing low yield is caused by environmental stress, and to correct the basic productivity value accordingly.