Method and device for determining stable yield threshold value of soil salinity in cotton root zone

By constructing a structured information database and using interpretable machine learning methods, the problem of cross-regional and cross-mode migration of the stable yield threshold of cotton root zone soil salinity was solved, and the standardized extraction and uncertainty quantification of the stable yield threshold were realized, thereby improving the scientific nature and engineering application of water and salt management.

CN121981602APending Publication Date: 2026-05-05CHINA AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA AGRI UNIV
Filing Date
2026-01-19
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing technologies for determining stable yield thresholds for cotton root zone soil salinity suffer from poor cross-regional and cross-planting pattern transferability, difficulty in quantifying uncertainty, and uninterpretable threshold results. In particular, it is difficult to achieve systematic and interpretable application of stable yield thresholds in multiple cotton-growing areas and under multiple irrigation patterns.

Method used

By constructing a structured information database of multi-source literature data, standardizing the yield calculation criteria and labeling the types, depths, and time points of salinity indicators, and combining crop models to simulate the dynamic process of salinity, interpretable machine learning methods are used to map stable yield thresholds and quantify uncertainties, outputting regionalized thresholds and their uncertainty ranges.

Benefits of technology

This study enabled the standardized extraction and reliable extrapolation across regions and models of the soil salinity threshold for stable cotton yield, improving the scientific nature and engineering application capabilities of water and salt management decisions, and quantifying the uncertainty and strength of evidence for the threshold.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121981602A_ABST
    Figure CN121981602A_ABST
Patent Text Reader

Abstract

The invention provides a cotton root zone soil salinity stable yield threshold determination method and device. Belongs to the technical field of agricultural water and soil resource management and the like, and aims at the problems of scattered threshold evidence, different calibers, poor mobility, uncertainties, difficulty in quantification and the like in the prior art, a structured salinity-yield database is constructed through literature integration, and a site scale stable yield threshold is extracted through index standardization and relative yield unification; a comparable mechanism feature is generated in combination with a crop model, and uncertainty is quantified through lightweight sampling; and training a threshold mapping model by utilizing interpretable machine learning to realize cross-cotton-area / cross-mode threshold extrapolation. According to the method, the regional stable yield salinity threshold value, the uncertainty interval and the evidence intensity can be output, the threshold value migration problem caused by the difference of root zone water and salt processes in different modes such as drip irrigation film mulching and non-drip irrigation is solved, and a scientific and explainable decision basis is provided for nationwide regional and different-mode cotton salinization regulation and control.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of agricultural water and soil resource management, salinized farmland regulation and crop yield stability threshold identification, and particularly to a method and apparatus for determining the yield stability threshold of cotton root zone soil salinity. Background Technology

[0002] Cotton, an important economic crop in my country, is widely cultivated in arid and semi-arid irrigated areas such as the Northwest Inland Region. These areas experience scarce rainfall and high evaporation and transpiration rates, making agricultural production highly dependent on irrigation. Due to differences in groundwater depth, soil texture, and irrigation systems, salt tends to accumulate and migrate seasonally in the crop root zone, significantly limiting transpiration, biomass formation, and yield stability during key cotton growth stages (such as flowering and boll formation). Therefore, rationally determining the salt-induced yield stability threshold for cotton in different regions and under different management models is a crucial fundamental issue in ensuring stable yields and improving efficiency in saline-alkali irrigated areas, as well as in irrigation management decisions.

[0003] Current engineering and agronomic research typically uses the electrical conductivity (ECe) of saturated soil paste extract in the rhizosphere to characterize salt stress levels, and fits the yield-salt relationship using a "threshold-yield reduction rate" response relationship (such as a piecewise linear function or a sigmoid function) to determine the initial threshold of salt stress and yield reduction parameters. In crop model systems, salt stress is usually reflected by the stress coefficient Ks and the corresponding ECe threshold parameter, indicating its inhibitory effect on transpiration, dry matter accumulation, and yield formation processes. In recent years, with the development of remote sensing technology and machine learning methods, research and patent practices have emerged that utilize multi-source data for soil salinity retrieval, yield prediction, or irrigation management optimization, aiming to improve spatial coverage and prediction accuracy.

[0004] However, while the aforementioned existing technologies have certain applicability under localized studies or single-management conditions, they still have significant limitations in forming cotton salinity yield stability thresholds at the "regional scale, cotton-specific area, or planting pattern" that can be used for engineering decision-making. First, different studies show inconsistencies in soil depth, sampling time (pre-sowing, critical growth period, or end of season), salinity index types, and conversion paths. Yield caliber, variety background, and water and fertilizer management conditions also vary significantly. Without a unified yield stability criterion and data traceability mechanism, directly integrating literature data can easily misinterpret differences in measurement methods as regional differences, leading to a systematic shift in threshold estimation. Second, different irrigation and planting patterns significantly alter water and salt transport processes. For example, drip irrigation with mulch and furrow or raised bed irrigation present fundamental differences in root zone wetting body morphology, leaching pathways, and spatial distribution of salt. This can cause the degree of salt exposure during critical growth periods to decouple from the end-of-season ECe characterization results. The same ECe value does not have equivalent meaning for yield stability under different management patterns, resulting in a significant decrease in the cross-pattern mobility of the threshold. Furthermore, traditional piecewise regression or empirical threshold fitting methods are highly sensitive to sample structure and salinity gradient coverage. Threshold points are prone to fluctuation under conditions of sparse samples or high noise levels, and they typically lack confidence intervals, evidence strength assessments, and uncertainty propagation mechanisms, making them unsuitable for engineering and regional applications. Meanwhile, while purely data-driven methods have advantages in salinity or yield prediction accuracy, they often use error minimization as the objective function, lacking intrinsic constraints on the physical semantics of the threshold, consistency of yield stabilization criteria, and cross-regional extrapolation boundaries. This makes it difficult to directly output salinity yield stabilization threshold results with clear interpretability and transferability.

[0005] Against this backdrop, providing an alternative technological approach has a clear practical motivation and application need. Field salinity threshold experiments are costly and time-consuming, making them difficult to conduct systematically in multiple cotton-growing areas with diverse irrigation and planting patterns. Meanwhile, agricultural management practices urgently require salinity yield stability threshold results that can reflect regional and management pattern differences and possess interpretability and uncertainty quantification capabilities. By constructing a traceable database of literature evidence, standardizing yield definitions across different studies with unified relative yield stability criteria, and introducing crop models to transform complex water-salt processes into comparable mechanistic characteristic parameters, combined with interpretable machine learning methods to characterize the response of yield stability thresholds to changes in regional environment and management patterns, this approach is expected to overcome the problems of non-transferable thresholds, inconsistent evidence, and difficulty in quantifying uncertainty in existing technologies. However, its final technological effect exhibits significant unpredictability. Summary of the Invention

[0006] The purpose of this invention is to provide a standardized extraction and regional extrapolation method and apparatus for stable yield threshold of cotton root zone soil salinity, which is used to output stable yield threshold and its uncertainty range in different cotton areas and under different planting / irrigation modes, and to provide evidence strength indicators.

[0007] Therefore, the first objective of this invention is to provide a method for determining the stable yield threshold of soil salinity in the cotton root zone.

[0008] Another objective of this invention is to provide a device for determining the stable yield threshold of cotton root zone soil salinity.

[0009] The third objective of this invention is to provide a computer device.

[0010] A fourth objective of this invention is to provide a non-transitory computer-readable storage medium.

[0011] To achieve the above objectives, a first aspect of the present invention provides a method for determining the stable yield threshold of cotton root zone soil salinity, comprising: S10. Construct a structured cotton salinity-yield information database containing multi-source literature data. Eliminate the impact of data heterogeneity by unifying the yield calculation caliber and labeling the salinity index type, measurement depth, time point and management mode. S20 extracts stable production salinity thresholds at the site-year scale based on a unified stable production criterion, and adds quality grade labels according to the salinity gradient coverage in the literature. S30 uses crop models to simulate the dynamic process of salinity in the root zone, extracts mechanistic features with physical semantics, and quantifies the uncertainty caused by the lack of key input parameters. S40 employs an interpretable machine learning model to map regional, pattern, soil-climate covariates, and mechanistic characteristics to stable yield thresholds, outputs regionalized thresholds and their uncertainty ranges, and analyzes feature contributions to support cross-regional extrapolation.

[0012] In one embodiment of the present invention, S10 includes: S101, collect and record literature metadata, experimental site information, experimental year information, planting / irrigation pattern information, soil and climate background information, root zone salinity index information and yield information, and establish a structured mapping relationship between site-year-treatment; S102, define SiteYearID to uniquely identify a site-year, define TreatmentID to uniquely identify different processing within the same SiteYear, and record the data source method and extraction error level.

[0013] In one embodiment of the present invention, S20 includes: S201. When other salt conductivity indices are given in the literature, they are converted to ECe according to their measurement conditions and conversion relationships. If a reliable conversion is not possible, the original index is retained and marked as non-ECe, and used as a covariate or removed during model training. S202, using formula Calculate the relative output and set the stable output criterion as Y. r ≥0.90.

[0014] In one embodiment of the present invention, S30 includes: S301, run the crop model to obtain the dynamic ECroot(t) of root zone salinity and the salt stress coefficient Ks(t), and calculate the stage salt stress integral:

[0015] As a mechanistic feature; S302, Extract the number of days of exposure exceeding the threshold:

[0016] and root zone salt peak:

[0017] As a characteristic of salt exposure.

[0018] In one embodiment of the present invention, S40 includes: S401, construct the input feature vector X, wherein the feature vector X includes at least cotton district, irrigation / planting pattern, soil texture, climate background, drought index and mechanism characteristics; S402 employs a cross-document validation strategy to assess extrapolation capability, preferentially selects leave-one-document cross-validation, and outputs model error statistics.

[0019] To achieve the above objectives, a second aspect of the present invention provides a device for determining the stable yield threshold of cotton root zone soil salinity, comprising: The structured information database construction module is used to build a structured cotton salinity-yield information database containing multi-source literature data. By standardizing the yield calculation method and labeling the salinity index type, measurement depth, time point and management mode, the impact of data heterogeneity is eliminated. The stable production threshold extraction and quality labeling module is used to extract the stable production salinity threshold at the site-year scale based on a unified stable production criterion, and to add quality grade labels according to the salinity gradient coverage in the literature. The crop model simulation and uncertainty quantification module is used to simulate the dynamic process of salinity in the root zone using crop models, extract mechanistic features with physical semantics, and quantify the uncertainty caused by the lack of key input parameters. The machine learning mapping and threshold output module is used to map regional, pattern, soil and climate covariates and mechanistic features to stable yield thresholds using an interpretable machine learning model, output regionalized thresholds and their uncertainty intervals, and analyze feature contributions to support cross-regional extrapolation.

[0020] The present invention discloses a method and apparatus for determining the stable yield threshold of cotton root zone soil salinity, which can unify the caliber of multi-source heterogeneous literature data, integrate crop model mechanism characteristics and interpretable machine learning, realize the standardized extraction of stable yield threshold of cotton root zone soil salinity and reliable extrapolation across regions and models, quantify the uncertainty of the threshold and the strength of evidence, and improve the scientific nature and engineering application capability of water and salt management decisions.

[0021] To achieve the above objectives, a third aspect of this application provides a computer device comprising a processor and a memory; wherein the processor runs a program corresponding to the executable program code stored in the memory, for implementing a method for determining the stable yield threshold of cotton root zone soil salinity as described in the first aspect embodiment.

[0022] To achieve the above objectives, the fourth aspect of this application proposes a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a method for determining the stable yield threshold of cotton root zone soil salinity as described in the first aspect embodiment.

[0023] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0024] Figure 1 This is a flowchart of a method for determining the stable yield threshold of cotton root zone soil salinity according to an embodiment of the present invention; Figure 2 This is a schematic diagram of a method for determining the stable yield threshold of cotton salinity according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the module structure of the cotton salinity stable yield threshold determination system according to an embodiment of the present invention; Figure 4 This is a schematic diagram of mechanism feature extraction according to an embodiment of the present invention; Figure 5 This is a flowchart of uncertainty propagation and threshold interval output according to an embodiment of the present invention; Figure 6 This is a schematic diagram of a regionalized threshold output strategy according to an embodiment of the present invention; Figure 7 This is a structural diagram of a cotton root zone soil salinity stable yield threshold determination device according to an embodiment of the present invention; Figure 8 It is a computer device according to an embodiment of the present invention. Detailed Implementation

[0025] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0026] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0027] The following describes, with reference to the accompanying drawings, a method and apparatus for determining the stable yield threshold of cotton root zone soil salinity according to an embodiment of the present invention.

[0028] Example 1 Figure 1 This is a flowchart of a method for determining the stable yield threshold of cotton root zone soil salinity according to an embodiment of the present invention, as shown below. Figure 1 As shown, it includes: S10. Construct a structured cotton salinity-yield information database containing multi-source literature data. Eliminate the impact of data heterogeneity by standardizing the yield calculation method and labeling the salinity index type, measurement depth, time point and management mode.

[0029] Specifically, the steps of this invention involve unifying the yield calculation criteria and systematically labeling the salt content index type, measurement depth, time point, and management mode to construct a structured cotton salt content-yield information database, thereby eliminating threshold estimation bias caused by data heterogeneity.

[0030] Further, this step first standardizes the salinity indicators used in different literature. Saturated soil extract conductivity (ECe) is preferred as the uniform indicator. If other conductivity indicators are provided in the literature (such as soil extract conductivity EC1:5 or soil aqueous solution conductivity ECw), standardization is performed based on their measurement conditions and conversion paths. For indicators that cannot be reliably converted, the original data are retained and labeled as "non-ECe," and treated as covariates in subsequent modeling or removed according to data quality rules.

[0031] Secondly, this step involves structured labeling of the depth and time points for salinity measurements. Measurement depths are typically within the root zone of 0–40 cm or 0–60 cm, and time points include pre-sowing, critical growth stages (such as flowering-boll formation), and the end of the season. When multi-layered salinity data is available, root zone weighted salinity can be further calculated to more accurately reflect the actual stress levels on the crop. Based on a unified salinity index and time scale, this step introduces relative yield (Y). rAs a criterion for stabilizing production, its calculation formula is as follows:

[0032] in For the output of a certain process, This refers to the maximum output across all processes at the same site within a given year. The stable production criterion is set as follows: A stable yield is defined as a relative yield of at least 90%. This criterion has clear physiological significance and can effectively eliminate yield differences caused by non-salinity factors such as variety and water and fertilizer systems, thereby improving the comparability and transferability of threshold estimation.

[0033] Furthermore, this step is applicable to integrating multi-source literature data to construct a national-scale cotton salinity-yield database, especially suitable for situations where data calibers differ under different geographical and management backgrounds, such as the Northwest Inland Cotton Region and the Yellow River Basin Cotton Region. Through unified standardization and metadata annotation, a high-quality, traceable data foundation can be provided for subsequent threshold extraction, crop model simulation, and machine learning modeling.

[0034] Furthermore, this step reduces the interference of data heterogeneity on threshold estimation and improves the fusion and interpretability of data from different studies. By using a unified stable yield criterion and a structured annotation system, consistent input variables are provided for subsequent steps involving crop model-based mechanistic feature extraction and threshold mapping model training. This is a crucial preliminary step for achieving stable yield salinity threshold extrapolation at the national, cotton-growing region, and model-specific levels.

[0035] Furthermore, S10 includes: S101, collect and record literature metadata, experimental site information, experimental year information, planting / irrigation pattern information, soil and climate background information, root zone salinity index information, and yield information, and establish a structured mapping relationship between site, year, and treatment.

[0036] Specifically, the aim is to construct a structured and traceable database of cotton salinity and yield data, providing high-quality data support for subsequent extraction of salinity-based yield stability thresholds and regional extrapolation. This step ensures the comparability and consistency of data across multiple dimensions by systematically collecting and recording experimental data.

[0037] Furthermore, this step first extracts and structures the metadata of publicly available literature, including title, author, publication year, and journal name, to support the traceability of data sources. Experimental site information must include the site name, latitude and longitude coordinates (or administrative division down to the city / county level), altitude (optional), and the cotton-growing region it belongs to (e.g., Northwest Inland Cotton Region, Yellow River Basin Cotton Region, etc.) to achieve spatial scale division. Experimental year information includes the year, sowing, emergence, and harvest dates, or an estimate of the growth period length; if information is missing, the level of missing information must be indicated. Planting / irrigation pattern information covers categories such as drip irrigation with mulch, drip irrigation without mulch, furrow irrigation, and raised bed irrigation, and records the number of irrigations, the amount of water per irrigation, and the irrigation regime to reflect the impact of management patterns on water and salt processes.

[0038] Furthermore, soil and climate background information, including soil layer thickness, texture (sand, silt, and clay ratio), field water holding capacity, wilting coefficient, and background salinity, are used as soil profile parameters for constructing crop models. For root zone salinity indicators, saturated soil electrical conductivity (ECe) is preferred, and the measurement depth, time point (pre-sowing, critical period, end of season), and measurement method are recorded to ensure consistent salinity characterization. Yield information must be labeled as seed cotton or lint cotton, and the treatment number and treatment description (such as salinity gradient source, irrigation water salinity, and salt discharge conditions) must be recorded to ensure consistent yield definitions.

[0039] Furthermore, this step defines a structured mapping relationship between site, year, and process, establishing SiteYearID and TreatmentID as data primary keys to ensure that each process has a unique identifier within a specific SiteYear. Simultaneously, it records the data source method (e.g., tables, text, digitized figures) and extracts error levels to quantify data quality.

[0040] Furthermore, this step requires all data fields to have clearly defined units and standards, such as ECe. It indicates that the salinity of irrigation water is... The length of the reproductive period is expressed in days. Furthermore, data quality grades (such as A, B, C) are used to identify the number of stable-yield points, salinity gradient coverage, and data completeness, providing a basis for subsequent uncertainty analysis.

[0041] Furthermore, this step is applicable to cotton salinization management studies across multiple regions, models, and literature sources. By unifying the data structure and metadata recording, experimental data from different sources, measurement methods, and management conditions can be effectively integrated, providing a data foundation for constructing stable yield salinity thresholds at a national scale.

[0042] Furthermore, this step significantly improves the comparability and traceability of data through structured data mapping and standardized metadata recording, providing reliable data support for subsequent threshold extraction, model training, and regionalized output. It is a key prerequisite for realizing cross-regional and cross-mode salinity stability threshold fusion analysis.

[0043] S102, define SiteYearID to uniquely identify a site-year, define TreatmentID to uniquely identify different processing within the same SiteYear, and record the data source method and extraction error level.

[0044] Specifically, this step aims to construct a basic data information database on cotton salinity and yield, and to achieve structured management and unique identification of the data by defining SiteYearID and TreatmentID. Further, SiteYearID is used to uniquely identify the "site-year" combination, ensuring that experimental data from the same location in different years can be accurately distinguished and correlated across different literature sources. TreatmentID is used to uniquely identify different treatment units within the same SiteYear, such as experimental treatments under different salinity gradients, irrigation methods, or water and salt management measures, thereby supporting subsequent multi-treatment comparisons and threshold extraction at the site-year scale.

[0045] Further, this step first collects and structures the literature metadata, including the title, author, publication year, and source journal, to ensure data traceability. Experimental site information includes site name, latitude and longitude coordinates (or administrative division down to the city / county level), cotton-growing region (e.g., Northwest Inland Cotton Region, Yellow River Basin Cotton Region, etc.), and optional altitude information. Planting / irrigation pattern information covers categories such as drip irrigation with mulch, drip irrigation without mulch, furrow irrigation, and raised bed irrigation, and records the number of irrigations, the amount of water per irrigation, and the irrigation sequence. Soil background information includes soil thickness, texture composition (sand, silt, clay ratio), field water holding capacity, wilting coefficient, etc., to support crop model input.

[0046] Furthermore, the data source methods need to be recorded in detail, such as from literature tables, text descriptions, or digitized figures, and the extraction error levels should be labeled (e.g., high, medium, low) to reflect the potential impact of data quality on subsequent analysis. Further, this step, by defining SiteYearID and TreatmentID, provides a foundation for data association and processing tracking for subsequent steps (such as relative yield calculation, stable yield threshold extraction, and crop model input construction), ensuring accurate identification of data units and control of data consistency and comparability during the integration of multi-source heterogeneous data. This step plays a core role in data standardization and structuring in the entire technical solution, providing high-quality, interpretable data support for subsequent model training and regional extrapolation.

[0047] S20 extracts stable production salinity thresholds at the site-year scale based on a unified stable production criterion, and adds quality grade labels according to the salinity gradient coverage in the literature.

[0048] Specifically, this step, "extracting stable salinity thresholds at the site-year scale based on a unified stable yield criterion, and adding quality grade labels according to the salinity gradient coverage in the literature," is technically based on a unified stable yield criterion. By screening salinity treatments that meet the criteria at the site-year scale, a stable salinity threshold is extracted, and a quality grade label is added based on the salinity gradient coverage to ensure the comparability and reliability of the threshold.

[0049] Furthermore, the relative output of processing within each SiteYear is first calculated. The calculation formula is:

[0050] in For the output of a certain process, This represents the maximum output across all processes within that Site Year. At that time, the treatment was considered to be within the range of stable production. Subsequently, the corresponding set of salinity indicators was extracted from all stable production treatments. And define the salinity threshold for stable production. The maximum value in the set:

[0051] If the number of stable production points is insufficient or the salinity gradient coverage is inadequate, interpolation or quantile approximation strategies should be used for estimation. And attach a quality grade label, such as Indicates a sufficient gradient and a stable production point. , Indicates the gradient general or steady-production point , This indicates that interpolation is required or there is a significant risk of ambiguity.

[0052] Furthermore, this step is applicable to integrating multi-source literature data and unifying the yield and salinity indices from different studies, especially under different irrigation modes (such as drip irrigation with mulch, furrow irrigation, etc.) and soil background conditions, ensuring the comparability and transferability of the extracted stable yield salinity thresholds. By adding quality grade labels, data quality and sources of uncertainty can be further identified, providing a reliable basis for subsequent model training and regional extrapolation.

[0053] Furthermore, this step, through standardized processing and quality control mechanisms, effectively alleviates the threshold estimation bias caused by inconsistent data caliber and incomplete sample structure in the literature, laying a data foundation for constructing a stable yield salinity threshold system at the national scale and by cotton region / model.

[0054] Furthermore, S20 includes: S201. When other salt conductivity indices are given in the literature, they are converted to ECe according to their measurement conditions and conversion relationships. If a reliable conversion is not possible, the original index is retained and marked as non-ECe, and used as a covariate or removed during model training.

[0055] Specifically, when the salinity index provided in the literature is not the standard root zone saturated extractable conductivity (ECe), it needs to be converted according to its measurement conditions and known conversion relationships to achieve uniformity of the salinity index. The technical implementation of this step is based on the systematic identification and standardization of salinity measurement methods, depths, and time scales, ensuring that data from different sources are compared and modeled under a unified physical dimension.

[0056] Furthermore, if the salinity index provided in the literature is electrical conductivity, it needs to be converted according to its measurement conditions (such as soil moisture content, extract ratio, measurement time point, etc.) and the definition of ECe. For example, EC1:1 refers to the electrical conductivity after soil and water are extracted at a 1:1 volume ratio, while ECe refers to the electrical conductivity extracted at saturated moisture content. During the conversion process, the relevant procedures of USSL saturated paste extraction or FAO / ASCE should be consulted to ensure the scientific validity and traceability of the conversion path.

[0057] Furthermore, when the literature does not provide sufficient conversion evidence or its measurement method lacks a clear conversion relationship with ECe, the indicator will be retained and labeled as "non-ECe". During the model training phase, such indicators can be used as covariates in modeling or removed according to data quality rules to avoid introducing systematic bias. For example, if a literature only provides ECw but does not specify the extract concentration or measurement time point, it cannot be reliably converted to ECe. In this case, it should be marked as "non-ECe" and its original measurement conditions should be recorded.

[0058] Furthermore, this step plays a crucial role in model building. By standardizing salinity indicators from different literature to ECe or retaining them as "non-ECe" and labeling them, threshold misjudgments caused by differences in measurement methods can be effectively reduced, improving the comparability of data across different site-year scales. Simultaneously, this standardization process provides a unified input basis for extracting the stable-yield salinity threshold in subsequent step S30, ensuring relative yield... The calculations are performed under a consistent salinity index system, thereby improving the accuracy and interpretability of the model predictions.

[0059] S202, using formula Calculate the relative yield and set the stable yield criterion as Yr≥0.90.

[0060] This step aims to address inconsistencies in salinity indices, measurement depths, time scales, and yield calibers across different literature sources. Through standardized processing, a comparable relative yield (Yr) index is constructed, providing a consistent input basis for subsequent threshold extraction and model training. In some implementations, this step first standardizes the salinity indices. When the literature directly provides root zone soil saturated paste conductivity (ECe), ECe is directly used as the unified salinity index. If the literature provides other salinity conductivity indices (such as ECw, ECe1:5, etc.), they are converted to ECe based on their measurement conditions and known conversion relationships. If reliable conversion is not possible, the original index is retained and labeled as non-ECe, and treated as a covariate in subsequent model training or removed according to quality rules.

[0061] Secondly, the depth and time scale of the salinity measurements are labeled. If the literature provides multi-layer salinity data, root zone weighted salinity can be further calculated to more accurately reflect the actual salinity stress experienced by the root system. If only a single depth or time point is provided, it is retained as metadata fields for subsequent weighting or filtering in the model.

[0062] In calculating relative yield, this step uses the following formula: ,in, The measured yield of a certain treatment. This represents the maximum yield among all treatments at that site within a given year. This formula normalizes the yields of different treatments to relative values, eliminating absolute yield deviations caused by differences in variety, water and fertilizer systems, etc. Furthermore, a stable yield criterion is set as follows: In other words, a stable production status is defined as when the relative yield is not lower than 90%. This criterion conforms to the general standards for agricultural production stability assessment and has strong interpretability and operability.

[0063] Furthermore, this step is typically completed during the data preprocessing stage and is applicable to the integration and cleaning of multi-source literature data. Its technical effect lies in significantly improving the comparability of data across different studies, providing unified and traceable input variables for subsequent threshold extraction and model training, thereby enhancing the applicability and stability of the entire system under cross-regional and cross-modal conditions.

[0064] S30 uses crop models to simulate the dynamic process of salinity in the root zone, extracts mechanistic features with physical semantics, and quantifies the uncertainty caused by the lack of key input parameters.

[0065] Specifically, this invention simulates the dynamic changes in salinity in the root zone of cotton by running a crop model and extracting mechanistic features with clear physical semantics to achieve a quantitative description of the salt stress process. This step is a key link connecting literature data with interpretable machine learning models. Its technical implementation is based on the crop model's ability to simulate water and salt transport processes, as well as time-series integration and feature extraction methods for salt stress intensity.

[0066] Furthermore, the crop model needs to have a salt stress response module, capable of outputting a time series of root zone salt concentrations based on input soil salinity distribution, meteorological conditions, irrigation regimes, etc. and salt stress coefficient ,in 1 represents no stress, and 0 represents complete stress. Salt stress intensity. Defined as It is used to quantify the degree to which salt inhibits crop transpiration and biomass accumulation.

[0067] Furthermore, targeting different stages of childbearing (e.g., all-season ALL, critical period K, or S10–S40), calculate the salt stress integral. Its definition is:

[0068] in The time step for the model is typically set to 1 day. This integral characterizes the cumulative effect of salt stress during a specific reproductive stage, has clear physiological significance, and can be used to compare stress processes under different management models.

[0069] In addition, the mechanistic characteristics of the extraction also include the root zone salt peak. Average salt content and the number of days of exposure exceeding the threshold Their definitions are as follows:

[0070]

[0071]

[0072] in It can be set to a preset reference value or a reference threshold based on the training set. For the stage The total number of time steps within the time frame. These features can reflect the intensity, duration, and peak impact of salinity stress from different dimensions, providing physically meaningful input variables for subsequent threshold mapping models.

[0073] Furthermore, in practical applications, this step is typically run under standardized SiteYear input conditions, combined with a minimum viable input package, to ensure the comparability of the model across different literature datasets. Simultaneously, through uncertainty quantification, the mechanistic features output by this step will include quantile information, thus providing a robust input basis for threshold prediction.

[0074] Furthermore, this step transforms the salt stress process into comparable mechanistic characteristics through crop models, effectively overcoming the heterogeneity issues in measurement methods, time scales, and management models across different literature. This provides physically interpretable input variables for subsequent threshold mapping models, significantly enhancing the model's generalization ability and decision support value.

[0075] Furthermore, S30 includes: S301, run the crop model to obtain the root zone salt dynamics ECroot(t) and salt stress coefficient Ks(t), and calculate the stage salt stress integral.

[0076] Specifically, this step is a crucial link in crop model simulation and mechanistic feature extraction. Its technical implementation principle is based on the simulation of the dynamic process of salinity in the root zone using a crop growth model. Combined with the calculation of the salt stress coefficient, it further extracts physically meaningful stress features for subsequent training of the threshold mapping model and output of regionalized stable yield thresholds. Furthermore, this step simulates the root zone salinity changes under different treatment conditions by calling the standardized crop model input package, outputting the time-series salinity conductivity. With salt stress coefficient ,in The range of values ​​is 1 indicates no salt stress, and 0 indicates complete stress.

[0077] First, based on the crop model input package, a crop growth model is run to simulate root zone salinity dynamics from sowing to harvest. The model output is provided daily or per-period. And calculate based on the mechanism of salt's inhibition of transpiration or biomass. Furthermore, according to the formula for defining salt stress intensity... The intensity of stress changes over time.

[0078] Furthermore, this step involves the calculation of several key features, including the phase salt stress integral. Its definition is:

[0079] in The time step is typically 1 day. Additionally, the root zone salinity peak is extracted. Average salt content and the number of days of exposure exceeding the threshold Its definition is:

[0080] in A preset reference salinity threshold is used to determine whether the salinity exceeds the critical value for stable production.

[0081] Furthermore, this step is typically run at a site-year scale to simulate salt stress processes for different irrigation / planting patterns (such as drip irrigation with mulch, furrow irrigation, etc.). The output is used as input features for a threshold mapping model to train an interpretable machine learning model, enabling cross-regional extrapolation of stable yield salt thresholds.

[0082] Furthermore, this step transforms the salinity process into physically meaningful stress characteristics through crop models, enhancing the comparability between different management models and data from different literature sources. Simultaneously, by combining uncertainty quantification methods (such as LHS sampling), quantiles of stress characteristics (such as P5, P50, and P95) can be output, providing a quantitative basis for the interval-based expression of stable yield thresholds and risk warnings, thereby enhancing the engineering application value of threshold products.

[0083] S302, Extract the number of days of exposure exceeding the threshold:

[0084] and root zone salt peak:

[0085] As a characteristic of salt exposure.

[0086] In this invention, the step "SaltDaysg" is used to quantify salt stress during a specific reproductive stage. The cumulative effect of salinity on cotton growth. Extraction of this feature depends on the output time series of the crop model and a preset salinity reference threshold. The technical implementation principle is as follows: Furthermore, during the operation of the crop model, the system outputs the daily root zone salinity conductivity. ,in Indicates the time step (usually 1 day). For any reproductive stage. The number of days of exposure exceeding the threshold is defined as the number of days during which the root zone salinity exceeds the reference threshold. The total number of days.

[0087] Furthermore, it is essential to first ensure that the time resolution of the crop model output is on a daily scale to guarantee... The continuity and comparability were then established. Subsequently, for each SiteYear and its typical treatment, key growth stages (such as flowering-boll formation) were extracted. The sequence is analyzed, and it is determined daily whether it exceeds [a certain limit]. This process can be implemented programmatically, for example, using Python's NumPy or Pandas libraries for vectorized computation, thus improving processing efficiency.

[0088] Furthermore, to address the uncertainty of input parameters, this invention employs Latin hypercube sampling (LHS) or Monte Carlo sampling methods to repeatedly run the crop model multiple times, thereby obtaining the probability distribution of SaltDaysg. The final output is in the form of quantiles (e.g., P5, P50, P95), where P50 serves as the main input for threshold mapping model training, and P5 and P95 are used to construct the uncertainty interval for the threshold, enhancing the robustness of the prediction results.

[0089] Furthermore, under drip irrigation and mulching conditions in the cotton-growing areas of Northwest China, SaltDaysg can be used to assess the cumulative duration of salt stress during critical growth periods, thereby assisting in the formulation of irrigation and drainage strategies. In addition, this feature, along with other mechanistic features (such as the salt stress integral SIg and the salt peak ECpeak), forms an input vector used to train an interpretable machine learning model, enabling regional extrapolation and uncertainty quantification of salt-stressed yield thresholds. By introducing SaltDaysg, this invention effectively improves the comparability and interpretability of salt stress processes under different management models, providing a scientific basis for agricultural water and salt regulation.

[0090] S40 employs an interpretable machine learning model to map regional, pattern, soil-climate covariates, and mechanistic characteristics to stable yield thresholds, outputs regionalized thresholds and their uncertainty ranges, and analyzes feature contributions to support cross-regional extrapolation.

[0091] Specifically, this invention employs an interpretable machine learning model to map regional, pattern, soil-climate covariates and mechanistic features to stable salinity thresholds, outputs regionalized thresholds and their uncertainty ranges, and analyzes feature contributions to support cross-regional extrapolation.

[0092] Furthermore, this step first extracts the stable salinity threshold. As a supervisory label And construct the input feature vector Its dimensions include geographical region, planting / irrigation pattern, soil background covariates (such as texture, initial salinity, field capacity, etc.), climate background covariates (such as drought degree, annual average precipitation, reference evapotranspiration, etc.), and extracted mechanistic features, such as salt stress integral. Number of days exposed to salt exceeding the threshold Root zone salt peak wait.

[0093] Furthermore, the model input features must meet interpretability requirements, such as using methods like SHAP or LIME to analyze feature contributions.

[0094] Furthermore, this step can be applied to predict stable yield salinity thresholds in different cotton-growing regions (such as the Northwest cotton-growing region and the Yellow River Basin cotton-growing region) under different planting / irrigation modes (such as drip irrigation with mulch and furrow irrigation). By analyzing the contribution of each feature to the threshold, key factors affecting threshold drift can be identified, providing a scientific basis for regional management.

[0095] Furthermore, this step, by introducing an interpretable machine learning model, not only improves the predictive accuracy of the stable production threshold but also enhances the model's transparency and traceability. Feature contribution analysis can reveal the mechanisms by which regional, pattern, and environmental variables influence the threshold, thereby supporting cross-regional extrapolation and management decisions.

[0096] Furthermore, S40 includes: S401, construct the input feature vector X, which includes at least cotton district, irrigation / planting pattern, soil texture, drought index and mechanism characteristics.

[0097] Specifically, constructing the input feature vector This is a crucial step in achieving regional extrapolation of the stable yield threshold for cotton root zone soil salinity. The core of this step lies in transforming heterogeneous literature data and crop model simulation results from multiple sources into a comparable and interpretable feature space, thereby providing structured input for subsequent machine learning modeling.

[0098] Furthermore, the input feature vector The construction of this data needs to encompass information from multiple dimensions. First, the geographical and management dimensions include cotton-growing regions and irrigation / planting patterns. Regions can be further subdivided into Northwest Inland Cotton-Growing Areas, Yellow River Basin Cotton-Growing Areas, Yangtze River Basin Cotton-Growing Areas, etc. Patterns include typical treatments such as drip irrigation with mulch, drip irrigation without mulch, furrow irrigation, and furrow irrigation. Second, soil and climate background information is used as covariates, such as soil texture (sand, silt, clay ratio), field water holding capacity, wilting coefficient, and drought indices (e.g., precipitation to evapotranspiration ratio), to characterize the heterogeneity of the regional water and salt environment.

[0099] Furthermore, the input feature vector Each dimension needs to define its value range and standardization method. For example, Region uses categorical coding (One-hot or LabelEncoding), and Mode is also a categorical variable; soil texture can be represented by triples (Sand, Silt, Clay), with units of percentage; drought index is usually expressed in a dimensionless form, such as... Where PET is potential evapotranspiration and P is annual precipitation. Mechanism characteristics are as follows: The cumulative effect of stress is quantified in the form of time integral, with units of dS·day. This is an integer variable representing the duration of the coercion, in days.

[0100] Furthermore, this feature vector will be used to train interpretable machine learning models (such as random forests, gradient boosting trees, SHAP interpretation methods, etc.) to achieve a stable salinity threshold at the SiteYear level. Extrapolation to larger scales (such as cotton-growing areas and management models) is required. Model inputs must encompass regional, model, soil, climate, and mechanistic characteristics to capture the drift patterns of salinity thresholds under different environmental and management conditions. For example, under drip irrigation and mulching conditions in the Northwest cotton-growing region, the model can identify the effects of soil texture and drought on salinity thresholds. This has a significant impact, thus providing a scientific basis for regional management.

[0101] Furthermore, the technical advantage of this step lies in unifying literature data and crop model output into an interpretable modeling framework through the construction of structured feature vectors, significantly improving the cross-regional transferability and model transparency of threshold prediction. Simultaneously, the uncertainty quantile information (such as P50, P5, and P95) contained in the feature vectors provides a basis for interval-based threshold output, enhancing risk control capabilities in engineering applications.

[0102] S402 employs a cross-document validation strategy to assess extrapolation capability, preferentially selects leave-one-document cross-validation, and outputs model error statistics.

[0103] Specifically, employing a cross-document validation strategy to evaluate the model's extrapolation ability, prioritizing the "leave-one-out cross-validation" method, and outputting model error statistics are crucial steps in ensuring that the constructed interpretable machine learning threshold mapping model possesses good generalization performance and transferability. This step is technically implemented based on a supervised learning framework, the core of which lies in using a rigorous validation mechanism to examine the model's predictive stability and consistency on untrained literature data, thereby evaluating its extrapolation ability under different research contexts.

[0104] Furthermore, the LOSO validation strategy is implemented as follows: each document is treated as an independent validation unit, and it is removed from the training set sequentially. The model built using the remaining documents is then used to predict the SiteYear threshold for that document. This strategy effectively avoids the risk of overfitting caused by multiple SiteYears sharing similar environments or management conditions within the same document, and is particularly suitable for datasets with diverse document sources and wide regional coverage. In each validation, the model prediction value... Compared with the true threshold Compare these values ​​and calculate error statistics, such as mean squared error (MSE), mean absolute error (MAE), and coefficient of determination R. 2To quantify the prediction accuracy and stability of the model.

[0105] Furthermore, the error statistics output includes not only numerical indicators but also the distribution of prediction residuals at the document level to identify prediction biases of the model under specific regional, pattern, or environmental conditions. For example, if the prediction error of a certain type of document (such as drip irrigation mulching experiments) is significantly higher than that of other types, it suggests that the model's extrapolation ability under this management method is limited, and targeted feature engineering or data augmentation strategies need to be introduced in subsequent model optimization.

[0106] The method for determining the stable yield threshold of cotton root zone soil salinity in this invention achieves standardized extraction and regional extrapolation of the stable yield threshold of cotton root zone soil salinity, improves the transferability and interpretability of the threshold under different cotton regions, planting patterns and environmental conditions, and quantifies the output threshold uncertainty and evidence strength to support precise water and salt management decisions.

[0107] Example 2 This invention aims to address the challenges posed by fragmented publicly available data, inconsistent definitions of salinity and yield indicators, missing key inputs, and varying root zone water and salt processes due to differences in management practices such as drip irrigation with mulch and non-drip irrigation. It establishes a unified yield criterion (relative yield Y) to address these issues. r (≥0.90) to achieve standardized extraction of the stable yield threshold of soil salinity in cotton root zone, interpretable extrapolation across cotton regions / modes, and quantitative output of threshold uncertainty interval and evidence strength.

[0108] This invention proposes a method for determining the stable yield threshold of cotton root zone soil salinity based on literature integration, crop model mechanistic characteristics, and interpretable machine learning. Figure 2 As shown, it includes the following steps: Step S1: Construct a basic data information database for cotton salinity and yield. The basic data information database includes at least literature metadata, experimental site information, experimental year information, planting / irrigation pattern information, soil and climate background information, root zone salinity index information, and yield information, and establishes a structured mapping relationship of "site-year-treatment". Step S2, Indicator Standardization and Relative Yield Construction: The salinity indicators, measurement depth / time scale, and yield caliber from different literature sources are standardized and labeled. The relative yield Y is calculated on a site-year scale. r And set the stable production criterion Y r ≥0.90; Step S3, Site-Year Stable Yield Salinity Threshold Extraction: Within each SiteYear, extract the stable yield salinity threshold S based on the stable yield criterion. 90 ; Step S4, construct the minimum workable input package for the crop model: construct the crop model input package for each SiteYear (and its typical treatment), the input package includes at least the meteorological sequence, soil profile parameters, cotton crop parameters, irrigation / planting management information and initial salinity information; Step S5, Crop Model Simulation and Mechanism Feature Extraction: Run the crop model to obtain root zone salinity dynamics and / or salt stress coefficient, and calculate comparable mechanism features such as stage salt stress integral, number of days of salt exposure exceeding the threshold, and peak / mean salt concentration in the root zone; Step S6, Lightweight Uncertainty Quantification: Set reasonable ranges for key input parameters that are missing or uncertain in literature and use Monte Carlo sampling or Latin hypercube sampling (LHS) to repeat the simulation to obtain quantile representations of mechanistic features and threshold prediction inputs. Step S7, Threshold Mapping Model Training and Interpretation: Using the threshold obtained in step S3 as the supervision label, and the region, pattern, soil climate covariates and the mechanistic features obtained in steps S5–S6 as inputs, train an interpretable machine learning threshold mapping model to achieve threshold extrapolation and output feature contributions. Step S8, Regionalized Threshold Output: For the target region (at least including the Northwest Cotton Region), input the corresponding variables under the specified irrigation / planting mode and optional growth stage conditions, output the regionalized stable yield salinity threshold and its uncertainty range, and can output evidence strength and data gap prompts.

[0109] Further, step S1, constructing a cotton salinity-yield basic data information database, includes: collecting and recording literature metadata, including literature title, author, publication year, source journal, etc.; collecting experimental site information, including site name, latitude and longitude coordinates (or administrative region to city / county level), altitude (optional), and cotton-producing region (e.g., Northwest Inland Cotton Region, Yellow River Basin Cotton Region, Yangtze River Basin Cotton Region, etc.); collecting experimental year and planting season information, including year, sowing / emergence / harvest date or inferred growth period length (indicating the missing level if missing); collecting planting / irrigation mode information, including drip irrigation with mulch, drip irrigation without mulch, furrow irrigation, furrow irrigation, etc., as well as irrigation frequency, single irrigation volume, or staged irrigation regime; and collecting soil background information, including soil layer. Thickness, texture or sandy-silty composition, field water holding capacity / wilt coefficient (optional), initial salinity or background salinity; collect root zone salinity index information and measurement caliber, including salinity index type (preferably ECe), measurement depth, measurement time point (pre-sowing / critical period / end of season) and measurement method; collect yield information and label the caliber (seed cotton / lint cotton), as well as treatment number, treatment description (salt gradient source, irrigation water salinity, salt discharge / leaching conditions, etc.); establish data key relationships: define SiteYearID to uniquely identify site-year, define TreatmentID to uniquely identify different treatments within the same SiteYear, and record the data source method (table, text, digitized graph) and extraction error level.

[0110] Furthermore, step S2, standardization of indicators and construction of relative yield, includes: Unification of salinity indicators: When the literature directly provides ECe, ECe is used as the unified indicator; when the literature provides other salinity conductivity indicators, they are converted to ECe based on their measurement conditions and conversion relationships. If a reliable conversion is not possible, the original indicator is retained and labeled as "non-ECe," and used as a covariate or removed during model training (controlled by quality rules); Depth / time scale labeling: The salinity measurement depth and time point are used as metadata fields in the subsequent model, or converted to root region weighted salinity (when multi-layer salinity information is available); Relative yield calculation: Within each SiteYear, the maximum yield of all treatments in that SiteYear is set as Y. max,site-year Calculate the relative output for each processing output Y:

[0111] The stable production criterion is set as follows: if and only if Y r A value ≥0.90 is considered stable production.

[0112] Furthermore, step S3, the extraction of the stable yield salinity threshold for the site and year, includes: the stable yield salinity threshold S. 90 Determination: Within each SiteYear, take the set of salinity indicators {ECe} corresponding to all treatments that meet the stable production criteria. i|Y r,i ≥0.90}, defined as: S 90 =max{ECe i |Y r,i ≥0.90} When the number of stable production points is insufficient or the salinity gradient coverage is insufficient, an interpolation or quantile approximation strategy is used to estimate the upper limit threshold of stable production, and a quality level label is attached to the threshold (e.g., A = sufficient gradient and stable production points ≥ 2; B = average gradient or stable production point = 1; C = interpolation required or there is obvious caliber risk).

[0113] Furthermore, step S4, which constructs the minimum operable input package for the crop model, includes: retrieving the site location, year, and management information of SiteYear from the basic data information database; constructing meteorological sequences (including at least precipitation and reference evapotranspiration or variables required to calculate reference evapotranspiration); soil profiles (layer thickness, texture, and hydraulic parameters); cotton crop parameters (using a general cotton parameter set and allowing phenological correction); irrigation regimes (method, frequency, timing, and single irrigation volume); and initial salinity information. When key inputs (such as pre-sowing soil moisture, root depth constraints, irrigation date accuracy, and salinity migration-related parameters) are missing from the literature, a reasonable range is set and marked with a "missing level" for sampling in step S6.

[0114] Furthermore, step S5, crop model simulation and mechanistic feature extraction, includes: running the crop model to obtain at least one salinity process output: root zone salinity dynamics EC root (t) and / or salt stress coefficient K s (t) (values ​​range from 0 to 1, where 1 indicates no salt stress); Salt stress intensity is defined as: I(t) = 1 - K s (t) For stage g (which can be the entire season ALL, critical period K, or S1–S4), calculate the salt stress integral:

[0115] Where Δt is the time step (Δt = 1 day when the time step is daily). Simultaneously extract salt exposure characteristics, including but not limited to: Root zone salt peak:

[0116] Mean salt concentration in the root zone:

[0117] Number of days of exposure exceeding the threshold:

[0118] in You can use a preset reference value or a reference threshold based on the training set.

[0119] Furthermore, step S6, lightweight uncertainty quantification, includes: setting an uncertain input parameter set Θ, wherein Θ includes at least one or more parameters related to initial water content, root depth constraint, irrigation time series offset / error, and salt migration / leaching efficiency; generating a sample set of Θ using Monte Carlo sampling or Latin hypercube sampling, repeatedly running the crop model for each SiteYear (and its typical treatment) to obtain the mechanistic feature distribution; outputting the mechanistic features (e.g., P5 / P50 / P95) in quantile form, using P50 as the main input for subsequent threshold mapping, and using P5 / P95 for threshold intervals and risk indications.

[0120] Furthermore, step S7, threshold mapping model training and interpretation, includes: using S90 obtained in step S3 as the supervision label y; constructing an input feature vector x, wherein x includes at least: cotton region, irrigation / planting mode, soil background covariates, climate background covariates, and mechanistic features obtained in steps S5–S6 (such as SI). g etc); training can explain machine learning regression models f( ), to satisfy:

[0121] in To predict the threshold, a cross-document validation strategy is used to evaluate the extrapolation capability, with "Leave One Document Cross-Validation (LOSO)" being the preferred option, and model error statistics are output. An interpretability analysis method is used to output feature contributions to explain the main controlling factors of threshold drift in different cotton-growing areas and under different modes.

[0122] Furthermore, step S8, regionalized threshold output, includes at least one method: Method A (typical condition method): Under the target area and specified mode / stage conditions, construct typical soil and climate inputs (take the median or dominant category of the sample in the area), first obtain typical mechanism features (P50 and P5 / P95) through steps S4–S6, and then input them into the threshold mapping model to obtain typical thresholds and intervals; Method B (regional distribution method): Sample from the empirical distribution of soil / climate / management conditions in the target area to generate a scenario set, obtain mechanism features respectively and predict thresholds in batches, output the median and quantile intervals of the threshold distribution, and output evidence density and gap hints.

[0123] In one embodiment of the present invention, a system for determining the stable yield threshold of cotton root zone soil salinity is proposed, such as... Figure 3 As shown, it includes an input layer, a processing layer, and an output layer.

[0124] In one embodiment of the present invention, the site-year stable salinity threshold S 90 Extraction such as Figure 4 As shown, a site and year (denoted as SiteYear-A) located in the Northwest cotton region of a certain literature are selected. This SiteYear includes several salinity treatments, and the ECe and yield Y of each treatment are recorded. The relative yield Y is calculated according to step S2. r When the criterion for stable production is Y r When ≥0.90, the ECe set corresponding to the stable production treatment is {2.0, 3.5} dS / m, then according to step S3, we get: S 90 =3.5 dS / m; and this threshold is labeled as the quality level (determined by the number of gradient points and the number of stable production points).

[0125] In one embodiment of the present invention, the generation and uncertainty quantification of crop model mechanistic features are used to construct the crop model input package for SiteYear-A, such as... Figure 5 As shown, this includes the annual meteorological sequence, soil profile parameters, drip irrigation with mulch film, and initial salinity information. Reasonable ranges were set for missing pre-sowing soil moisture and root depth constraints. K sets of parameters were generated using LHS sampling, and the crop model was repeatedly run. The model outputs the daily root zone salinity EC. root (t) and salt stress coefficient K s (t). Calculate the salt stress integral SI for the critical period according to step S5. g The number of days of exposure exceeding the threshold, etc., are used to summarize their P5 / P50 / P95 values ​​for subsequent mapping model input and interval-based output.

[0126] In one embodiment of the present invention, the threshold mapping model training and the threshold output for the northwest region are as follows: Figure 6 As shown, the summaries of S from multiple SiteYears 90 As supervisory labels, the input features include Region, Mode, soil texture and background salinity, drought index, and mechanistic features (SI). g SaltDays g (Salt peak, etc.). The threshold mapping model is trained and validated with leave-one-out literature. Under drip irrigation and mulching conditions in the Northwest cotton region, typical thresholds and intervals are output using method A, or the median and quantile intervals of the threshold distribution are output using method B, which involves sampling scenario sets. This forms a stable salinity threshold product that can be directly used for regional management.

[0127] This invention unifies the yield criteria across different literatures by using relative yield stability criteria, introduces crop model salt stress process features to enhance cross-model comparability, employs lightweight sampling to pass on the uncertainty of missing inputs, and achieves regional extrapolation and contribution analysis of thresholds through interpretable machine learning. This improves the transferability and interpretability of thresholds applied at the national scale, and provides threshold ranges and evidence strength to support engineering decision-making.

[0128] Example 3 To achieve the above embodiments, such as Figure 7 As shown, this embodiment also provides a cotton root zone soil salinity stable yield threshold determination device 10, including: The structured information database construction module 100 is used to construct a structured cotton salinity-yield information database containing multi-source literature data. By unifying the yield calculation caliber and labeling the salinity index type, measurement depth, time point and management mode, the impact of data heterogeneity is eliminated. The stable production threshold extraction and quality labeling module 200 is used to extract the stable production salinity threshold at the site-year scale based on a unified stable production criterion, and to add quality grade labels according to the salinity gradient coverage in the literature. The crop model simulation and uncertainty quantification module 300 is used to simulate the dynamic process of salinity in the root zone using crop models, extract mechanistic features with physical semantics, and quantify the uncertainty caused by the lack of key input parameters. The machine learning mapping and threshold output module 400 is used to map regional, pattern, soil and climate covariates and mechanistic features to stable yield thresholds using an interpretable machine learning model, output regionalized thresholds and their uncertainty intervals, and analyze feature contributions to support cross-regional extrapolation.

[0129] Furthermore, the structured information database construction module 100 is also used for: Collect and record literature metadata, experimental site information, experimental year information, planting / irrigation pattern information, soil and climate background information, root zone salinity index information, and yield information, and establish a structured mapping relationship between site, year, and treatment; Define SiteYearID to uniquely identify a site-year, define TreatmentID to uniquely identify different processing within the same SiteYear, and record the data source method and extraction error level.

[0130] Furthermore, the stable production threshold extraction and quality identification module 200 is also used for: When other salt conductivity indices are given in the literature, they are converted to ECe according to their measurement conditions and conversion relationships. If a reliable conversion is not possible, the original index is retained and marked as non-ECe, and used as a covariate or removed during model training. Using formula Calculate the relative output and set the stable output criterion as Y. r ≥0.90.

[0131] The present invention discloses a device for determining the stable yield threshold of cotton root zone soil salinity. This device can unify the caliber of multi-source heterogeneous literature data, integrate crop model mechanism characteristics and interpretable machine learning, realize the standardized extraction of stable yield threshold of cotton root zone soil salinity and reliable extrapolation across regions and models, quantify the threshold uncertainty and evidence strength, and improve the scientific nature and engineering application capability of water and salt management decisions.

[0132] Example 4 To implement the methods of the above embodiments, the present invention also provides a computer device, such as... Figure 8 As shown, the computer device 600 includes a memory 601 and a processor 602; wherein, the processor 602 reads the executable program code stored in the memory 601 to run a program corresponding to the executable program code, so as to implement the various steps of the method for determining the stable yield threshold of cotton root zone soil salinity described above.

[0133] Example 5 To implement the above embodiments, this application also proposes a non-transitory computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a method for determining the stable yield threshold of cotton root zone soil salinity as described in the foregoing embodiments.

[0134] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0135] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

Claims

1. A method for determining the stable yield threshold of soil salinity in the root zone of cotton, characterized in that, include: S10. Construct a structured cotton salinity-yield information database containing multi-source literature data. Eliminate the impact of data heterogeneity by unifying the yield calculation caliber and labeling the salinity index type, measurement depth, time point and management mode. S20 extracts stable production salinity thresholds at the site-year scale based on a unified stable production criterion, and adds quality grade labels according to the salinity gradient coverage in the literature. S30 uses crop models to simulate the dynamic process of salinity in the root zone, extracts mechanistic features with physical semantics, and quantifies the uncertainty caused by the lack of key input parameters. S40 employs an interpretable machine learning model to map regional, pattern, soil-climate covariates, and mechanistic characteristics to stable yield thresholds, outputs regionalized thresholds and their uncertainty ranges, and analyzes feature contributions to support cross-regional extrapolation.

2. The method as described in claim 1, characterized in that, S10 includes: S101, collect and record literature metadata, experimental site information, experimental year information, planting / irrigation pattern information, soil and climate background information, root zone salinity index information and yield information, and establish a structured mapping relationship between site-year-treatment; S102, define SiteYearID to uniquely identify a site-year, define TreatmentID to uniquely identify different processing within the same SiteYear, and record the data source method and extraction error level.

3. The method as described in claim 1, characterized in that, S20 includes: S201. When other salt conductivity indices are given in the literature, they are converted to ECe according to their measurement conditions and conversion relationships. If a reliable conversion is not possible, the original index is retained and marked as non-ECe, and used as a covariate or removed during model training. S202, using formula Calculate the relative output and set the stable output criterion as Y. r ≥0.

90.

4. The method as described in claim 1, characterized in that, S30 includes: S301, run the crop model to obtain the dynamic ECroot(t) of root zone salinity and the salt stress coefficient Ks(t), and calculate the stage salt stress integral: As a mechanistic feature; S302, Extract the number of days of exposure exceeding the threshold: and root zone salt peak: As a characteristic of salt exposure.

5. The method as described in claim 1, characterized in that, S40 includes: S401, construct the input feature vector X, wherein the feature vector X includes at least cotton district, irrigation / planting pattern, soil texture, drought index and mechanism characteristics; S402 employs a cross-document validation strategy to assess extrapolation capability, preferentially selects leave-one-document cross-validation, and outputs model error statistics.

6. A device for determining the stable yield threshold of cotton root zone soil salinity, characterized in that, include: The structured information database construction module is used to build a structured cotton salinity-yield information database containing multi-source literature data. By standardizing the yield calculation method and labeling the salinity index type, measurement depth, time point and management mode, the impact of data heterogeneity is eliminated. The stable production threshold extraction and quality labeling module is used to extract the stable production salinity threshold at the site-year scale based on a unified stable production criterion, and to add quality grade labels according to the salinity gradient coverage in the literature. The crop model simulation and uncertainty quantification module is used to simulate the dynamic process of salinity in the root zone using crop models, extract mechanistic features with physical semantics, and quantify the uncertainty caused by the lack of key input parameters. The machine learning mapping and threshold output module is used to map regional, pattern, soil and climate covariates and mechanistic features to stable yield thresholds using an interpretable machine learning model, output regionalized thresholds and their uncertainty intervals, and analyze feature contributions to support cross-regional extrapolation.

7. The apparatus as claimed in claim 6, characterized in that, The structured information database construction module is also used for: Collect and record literature metadata, experimental site information, experimental year information, planting / irrigation pattern information, soil and climate background information, root zone salinity index information, and yield information, and establish a structured mapping relationship between site, year, and treatment; Define SiteYearID to uniquely identify a site-year, define TreatmentID to uniquely identify different processing within the same SiteYear, and record the data source method and extraction error level.

8. The apparatus as claimed in claim 6, characterized in that, The stable production threshold extraction and quality identification module is also used for: When other salt conductivity indices are given in the literature, they are converted to ECe according to their measurement conditions and conversion relationships. If a reliable conversion is not possible, the original index is retained and marked as non-ECe, and used as a covariate or removed during model training. Using formula Calculate the relative output and set the stable output criterion as Y. r ≥0.

90.

9. A computer device, characterized in that, Including processor and memory; The processor reads executable program code stored in the memory to run a program corresponding to the executable program code, so as to implement the method for determining the stable yield threshold of cotton root zone soil salinity as described in any one of claims 1-5.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements a method for determining the stable yield threshold of cotton root zone soil salinity as described in any one of claims 1-5.