Tobacco leaf production decision making method and system fusing climate data with fertilization strategies

CN122840443APending Publication Date: 2026-09-29FUJIAN TOBACCO +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611361730.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-09-03
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

然而,不同田块因土壤、坡向及作物群体结构差异,作物对相同插值气候的生理响应敏感性存在明显的空间异质性,若不进行校正,所用气候因子便无法准确反映作物实际受到的气候胁迫,导致模型输入存在系统性偏差

Benefits of technology

1. 因利用偏最小二乘回归从田块多季节插值气候序列与产量、品质实测数据中提取标准化回归系数向量作为气候响应特征向量,并基于该向量之间的余弦相似度构造空间影响权重,对常规克里金插值初始气候的邻近田块偏差进行加权校正,从而生成符合作物实际生理感受的有效气温、有效降水和有效日照驱动因子序列;该过程以响应模式相似度替代固定地理距离,直接降低了复杂地形下因插值气候偏差引起的作物胁迫估计误差,为施肥决策提供了更高质量的空间化气候输入。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122840443A_ABST
    Figure CN122840443A_ABST
Patent Text Reader

Abstract

The present application belongs to the field of agricultural information technology, and particularly relates to a tobacco production decision-making method and system fusing climate data and fertilization strategies. The method extracts a climate response characteristic vector, corrects climate bias based on spatial similarity, generates effective climate driving factors, trains a Bayesian multi-output response surface model, jointly predicts yield and multi-dimensional quality, generates an optimal fertilization scheme using a multi-objective evolutionary algorithm, and updates the model based on measured data in a closed loop. The present application eliminates climate data prediction bias, significantly improves yield and quality prediction accuracy, realizes the coordinated optimization of yield and quality, and provides scientific decision-making support for precise tobacco production.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of agricultural information technology, specifically to a method and system for tobacco production decision-making that integrates climate data and fertilization strategies. Background Technology

[0002] In hilly and mountainous tobacco-growing areas, due to the uneven distribution of meteorological stations, spatial interpolation methods (such as Kriging) are typically used to estimate daily temperature, precipitation, and sunshine duration for each field, and these interpolated data are directly used as input to fertilization decision-making models. However, due to differences in soil, slope aspect, and crop canopy structure, the physiological response sensitivity of crops to the same interpolated climate varies significantly across different fields. Without correction, the climate factors used cannot accurately reflect the actual climate stress experienced by the crops, leading to systematic biases in the model input. Furthermore, existing fertilization effect models often employ static multiple regression or quadratic fertilizer effect equations, which are insufficient to effectively characterize the nonlinear responses of flue-cured tobacco yield and multidimensional quality indicators such as nicotine and total sugar to nitrogen, phosphorus, and potassium inputs. They also struggle to capture the complex interactions between climate conditions and basic soil fertility, thus limiting the predictive power and decision-making accuracy of models in hilly and mountainous environments with significant interannual climate fluctuations.

[0003] Furthermore, tobacco cultivation is a non-stationary system. Continuous cropping leads to gradual changes in soil fertility, and variety replacement, causing static models trained on historical data to gradually become inaccurate over time. Current technologies generally lack mechanisms for sequentially updating model parameters and crop climate-sensitive characteristics, and fail to effectively identify and filter abnormal observation data caused by extreme weather or sampling errors in field production. This results in single-season outliers potentially directly contaminating model updates or decision outputs, reducing the robustness of recommended fertilization schemes and their ability to be applied continuously across seasons. Therefore, how to transform physically interpolated climate into effective climate drivers under conditions of limited historical data and significant spatial heterogeneity, construct a joint decision-making model that adaptively optimizes yield and quality, and achieve robust evolution of the model and climate-sensitive patterns in non-stationary environments, is a pressing technical problem that needs to be solved. Summary of the Invention

[0004] In existing technologies, how to transform physically interpolated climate into effective climate driving factors under conditions of limited historical data and significant spatial heterogeneity, construct a joint decision-making model that can adaptively optimize yield and quality, and achieve robust evolution of the model and climate-sensitive patterns in non-stationary environments is a pressing technical problem that needs to be solved. To address this problem, this invention provides solutions in the following aspects.

[0005] In a first aspect, this invention provides a tobacco production decision-making method integrating climate data and fertilization strategies, employing the following technical solution: acquiring historical interpolated climate sequences and yield and quality data for each field; extracting climate response feature vectors for the fields using partial least squares regression; calculating the cosine similarity between the target field and neighboring fields based on the climate response feature vectors to construct spatial influence weights; using the spatial influence weights to correct deviations in the historical interpolated climate sequences, generating a historical effective climate driving factor sequence; using the statistics of the climate driving factor sequence, historical available soil nutrients, and yield and quality data as a training set to train a Bayesian multi-output response surface model; and generating the effective climate driving factor sequence for the current field in the same manner during the current production season. The statistical values ​​are extracted and input into the Bayesian multi-output response surface model along with the current available soil nutrients and the set candidate fertilization amounts. This jointly predicts yield and multidimensional quality indicators. Based on preset quality targets and weights, the quality coordination index is calculated. The optimal fertilization scheme is generated using a multi-objective evolutionary algorithm and the ideal point method. After each growing season, measured yield and quality data are collected. Outliers are removed after validity testing. The measured yield and quality data are added to the training set as new samples. The parameters of the Bayesian multi-output response surface model are updated by re-maximizing the logarithmic marginal likelihood. The currently used climate response feature vector is smoothly updated based on the similarity between the old and new vectors when the sample size condition is met. The updated results are used for decision-making in the next season.

[0006] Preferably, the historical interpolated climate series includes daily temperature estimation series, daily precipitation estimation series, and daily sunshine duration estimation series; yield is the yield per unit area of ​​tobacco leaves, and quality data includes nicotine content and total sugar content; the climate response feature vector of the field is extracted using partial least squares regression, which includes: using the daily temperature estimation series, daily precipitation estimation series, and daily sunshine duration estimation series as independent variables, and using the yield per unit area of ​​tobacco leaves, nicotine content, and total sugar content as multiple response dependent variables, and extracting a standardized regression coefficient vector as the climate response feature vector; when the historical interpolated climate series of a certain field and the yield and quality data are insufficient for a preset number of growing seasons, the average climate response feature vector of all fields in the region is used as the climate response feature vector of that field.

[0007] Preferably, constructing the spatial influence weight includes: taking the target field as the center, forming a set of neighboring fields within a preset radius; calculating the cosine similarity between the target field and each neighboring field in the set on the climate response feature vector; performing a non-negative truncation on the cosine similarity, and then normalizing the non-negative truncation cosine similarity by exponentiation with a preset sharpening factor to obtain the normalized spatial influence weight.

[0008] Preferably, the process of using spatial influence weights to correct the deviation of historical interpolated climate and generate a historical effective climate driving factor sequence includes: calculating the initial interpolated climate deviation of each neighboring field relative to the target field; using spatial influence weights to weight and sum the initial interpolated climate deviations to obtain correction factors; and superimposing the correction factors onto the initial interpolated climate of the target field to obtain a historical effective climate driving factor sequence; wherein, the historical effective climate driving factor sequence includes an effective temperature sequence, an effective precipitation sequence, and an effective sunshine sequence.

[0009] Preferably, the statistical measures for extracting the climate driving factor sequence include: dividing the historical effective climate driving factor sequence into the rooting period, the vigorous growth period and the maturity period according to the tobacco growth period; extracting the average temperature, cumulative precipitation and average sunshine in each period respectively, and combining them to form statistical measures.

[0010] Preferably, the Bayesian multi-output response surface model is a Bayesian Gaussian process regression model or a Bayesian multinomial regression model; the input variables of the Bayesian multi-output response surface model include statistics, soil available nitrogen content, soil available phosphorus content, soil available potassium content, nitrogen application rate, phosphorus application rate and potassium application rate, and the output variables include tobacco yield per unit area, nicotine content, total sugar content and potassium content.

[0011] Preferably, the quality coordination index is defined as follows: the square of the difference between the predicted value of each quality index in the multidimensional quality index and its corresponding enterprise preset quality target value is calculated; the square deviations are weighted and summed using the enterprise preset quality target weight coefficients; the weighted sum is added to a zero division protection constant and the reciprocal is taken to obtain the quality coordination index, so that the larger the value of the quality coordination index, the better the overall quality coordination.

[0012] Preferably, generating the optimal fertilization scheme using a multi-objective evolutionary algorithm and the ideal point method includes: using nitrogen, phosphorus, and potassium application rates as decision variables, within a preset fertilization rate constraint range, employing a fast non-dominated sorting genetic algorithm to simultaneously maximize the predicted yield and quality coordination index, and obtaining a Pareto front solution set; combining fertilizer cost constraints, using the ideal point method to screen out the compromise fertilization scheme with the optimal comprehensive distance from the Pareto front solution set, as the optimal fertilization scheme.

[0013] Preferably, the validity check includes: when the historical yield records are insufficient for a preset number of growing seasons, if the measured yield is lower than a first preset percentage of the lowest historical yield or higher than a second preset percentage of the highest historical yield, the sample corresponding to the measured yield is determined to be abnormal data; when the historical yield records do not meet the preset historical data accumulation requirements, the lowest and highest historical yields of all fields in the planting area where the field is located are used as reference boundaries for determination; if any item in the measured quality data exceeds the preset physiologically feasible limit for tobacco leaves, the corresponding sample is also determined to be abnormal data; the data determined to be abnormal are removed and not included in subsequent update calculations.

[0014] Secondly, this application provides a tobacco production decision-making system that integrates climate data and fertilization strategies, employing the following technical solution: The tobacco production decision-making system that integrates climate data and fertilization strategies includes: a processor and a memory; the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the above-mentioned method for tobacco production decision-making that integrates climate data and fertilization strategies is implemented.

[0015] The embodiments of the present invention have at least the following beneficial effects: 1. By using partial least squares regression to extract standardized regression coefficient vectors from multi-season interpolated climate sequences and measured yield and quality data of fields as climate response feature vectors, and constructing spatial influence weights based on the cosine similarity between these vectors, the bias of the initial climate in conventional Kriging interpolation of neighboring fields is weighted and corrected, thereby generating effective temperature, effective precipitation, and effective sunshine driving factor sequences that conform to the actual physiological experience of crops. This process replaces fixed geographical distance with response pattern similarity, directly reducing the crop stress estimation error caused by interpolation climate bias under complex terrain, and providing higher quality spatialized climate input for fertilization decisions.

[0016] 2. By constructing a Bayesian multi-output response surface model with effective climate-driven statistics, available soil nutrients, and candidate fertilization amounts as inputs, and yield and multi-dimensional quality indicators such as nicotine, total sugar, and potassium content as joint outputs, and combining it with a quality coordination index defined based on the inverse of the quality target deviation and a fast non-dominated sorting genetic algorithm, the Pareto front is automatically searched and a compromise solution is selected using the ideal point method. This allows the decision model to directly output the optimal quantitative fertilization scheme that takes into account both yield and multiple quality indicators from the continuous fertilization space, without the need for repeated manual weight setting and trial and error. It also effectively characterizes the nonlinear drift of the optimal fertilizer application rate caused by changes in climate factors.

[0017] 3. After each harvest, the measured yield and quality data are validated using a method that includes the ratio of historical yield extremes and the physiologically feasible limits of tobacco leaves. After removing outlier samples, Bayesian posterior updates are performed on the response surface model parameters to adaptively constrain parameter variation. Simultaneously, the cosine similarity between the old and new climate response feature vectors is calculated when the number of seasons with accumulated valid data reaches a preset threshold. The smoothing coefficient is adjusted based on the mutation detection results, and the old and new vectors are fused using exponential moving average. This combination of dual-channel smoothing and multi-level protection criteria enables the system to automatically isolate outlier data contamination and the impact of extreme climate in a single season while absorbing new season environmental information. This maintains the cross-seasonal consistency of recommended fertilization amounts, thereby reducing interannual fluctuations in tobacco leaf quality and enhancing decision robustness under unstable planting conditions. Attached Figure Description

[0018] Figure 1 The flowchart illustrates the steps of the tobacco production decision-making method and system that integrates climate data and fertilization strategies in this invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings. (Refer to...) Figure 1 The tobacco production decision-making method integrating climate data and fertilization strategies includes steps S1-S5, as detailed below: S1: Obtain historical interpolated climate sequences and yield and quality data for each field, and use partial least squares regression to extract the climate response feature vector of the field.

[0021] In hilly and mountainous tobacco-growing areas, the complex terrain and uneven distribution of meteorological stations lead to systematic discrepancies between the estimated daily temperature, precipitation, and sunshine duration at the field level obtained using commonly used spatial interpolation methods (such as Kriging) and the actual microclimate. Furthermore, differences in soil type, slope aspect, and tobacco canopy structure across different fields result in inconsistent physiological responses among tobacco plants, even under the same interpolated climatic values. This spatial heterogeneity in crop sensitivity to climatic factors across fields makes it difficult for physical interpolation data to directly reflect the actual climatic stress experienced by crops. To quantify these differences in sensitivity patterns into computable features, this step utilizes limited historical yield and quality data from each field to extract climatic response feature vectors that characterize the sensitivity of crops to climatic factors.

[0022] In this embodiment, the preset number of growing seasons is set to at least three complete flue-cured tobacco growing seasons to ensure the temporal representativeness and analytical reliability of historical data; in other embodiments, this preset number can be adaptively adjusted according to the actual application scenario or data acquisition conditions.

[0023] Specifically, for the target field, historical records of at least three complete flue-cured tobacco growing seasons were continuously collected. For each historical growing season, a daily average temperature estimation series spatially interpolated and assigned to the field was obtained. (Unit: °C) Daily precipitation estimation series (Unit: mm) and daily sunshine duration estimation series (Unit: h), where The date sequence is from the date of transplanting. This represents the actual number of days in the fertile period for that season. Number of days in the fertile period for each season. They may differ; a uniform maximum reproductive period of 30 days should be preset during processing. The default parameter is 130 days, which is based on experience. However, it can be adjusted by the implementer according to the specific implementation scenario. During the growth season, zeros are padded to the end of the sequence to bring it to a length of [length missing]. ;for During the growing season, before harvesting Data for the day. Simultaneously, the actual yield per unit area of ​​tobacco leaves was collected at the end of the growing season. (Unit: kg·hm) -2 Nicotine content (mass percentage, dimensionless) and total sugar content (Mass percentage, dimensionless). The three climate sequences processed for each season are pieced together end-to-end in the order of temperature, precipitation, and sunshine to form a sequence of length [length missing]. The independent variable vector; combining output and the two quality indicators to form a vector of length... The dependent variable vector. When the target field has accumulated Data for one valid growing season and When, the matrix of independent variables is obtained. and dependent variable matrix .

[0024] Based on the independent variable matrix and dependent variable matrix A multi-response prediction model was established using partial least squares regression (PLSR). PLSR is a well-known technique that obtains the regression relationship by extracting latent variables from the independent and dependent variable spaces and maximizing their covariance. The calculation first involves processing the independent variable matrix... and dependent variable matrix The mean centering and unit variance scaling are performed, and then the regression coefficient matrix is ​​estimated iteratively. , making Number of latent variables Instead of specifying beforehand, leave-one-out cross-validation is used to select the option that minimizes the sum of squared total predicted residuals of the dependent variable. The value is determined automatically by the algorithm, resulting in the regression coefficient matrix. Then, expand it column-wise into a one-dimensional vector. Then perform the following steps on the vector: Norm normalization; This is the standardized regression coefficient vector extracted from partial least squares regression. Divide by zero by the protection constant, and take the value of Climate response feature vector The calculation formula is as follows: In the formula, This is the desired field climate response feature vector, with dimensions of... Each component corresponds to the standardized sensitivity of a certain climatic factor (temperature, precipitation or sunshine) on the comprehensive impact of a certain growth sequence on yield and quality. The vector magnitude is 1, so only directional information is retained. This represents the Euclidean norm.

[0025] When the effective growing seasons of the target field When the number of climatic response eigenvectors is less than 3, the above partial least squares regression cannot be stably performed. In this case, all fields within the planting area of ​​the field that have records of at least 3 valid growing seasons are retrieved, and the element-wise arithmetic mean of their climatic response eigenvectors is calculated: In the formula, This refers to the set of fields within the region that meet the specified conditions. The number of plots in the set. For fields The extracted climate response feature vector. The vector is directly assigned to the target field as its initial climate response feature vector. Thus, each field has obtained a feature vector that reflects the comprehensive sensitivity pattern of its crops to climate factors. The directional differences of this vector directly correspond to the degree of similarity or dissimilarity of the climate response patterns among the fields.

[0026] S2: Based on the climate response feature vector, calculate the cosine similarity between the target field and neighboring fields to construct spatial influence weights. Use the spatial influence weights to correct the bias of historical interpolated climate series and generate a historical effective climate driving factor series.

[0027] After obtaining the climate response feature vectors for each field, the conventional spatial interpolation of field-level temperature, precipitation, and sunshine data exhibits systematic biases due to the influence of hilly and mountainous terrain. Furthermore, different fields show varying physiological responses to the same interpolated climate, making it impossible to accurately reflect the actual climate stress experienced by crops by directly using the interpolated data. Therefore, for each historical growing season of each field, it is necessary to spatially correct the historical interpolated climate using the extracted climate response feature vectors to generate a sequence of historically effective climate driving factors that can represent the actual experience of crops.

[0028] For a given historical growing season of target field i, first obtain the initial temperature sequence for that season, which is then assigned to target field i using conventional kriging interpolation. (Unit: °C) Initial precipitation sequence (Unit: mm) and initial sunshine duration sequence (Unit: h). With target field i as the center, a preset radius is defined. Neighboring fields Form a set , The empirical value of 3km is used as a preset parameter, but it can be adjusted by the implementer according to the specific implementation scenario. (Regarding the set...) Each neighboring plot Calculate its climate response characteristic vector With the feature vector of the target field Cosine similarity between them: In the formula, For vector dot product, For the Euclidean norm, The range of values ​​is Dimensionless. Because and All have been carried out in advance Norm normalization makes the above equation equivalent to an inner product operation. This similarity measure assesses the directional consistency of crop sensitivity patterns to climate factors between two fields.

[0029] Non-negative truncation of the cosine similarity yields the orientation consistency. To eliminate neighboring fields with conflicting response patterns Subsequently, a sharpening factor was introduced. Increase the influence weight of high-scoring neighborhoods. The empirical value of 2 is used; it is dimensionless and a preset parameter, but can be adjusted by the implementer according to the specific implementation scenario. Orientation consistency will be considered. Perform exponentiation and normalization to obtain the spatial influence weights. : In the formula, For orientation consistency, the denominator is the set. The sum of similarities of all sharpened fields, with spatial influence weights. satisfy And only if hour Dimensionless. This space influences the weights. Instead of fixed geographical distance, climate response pattern similarity is used to concentrate the correction contribution on neighboring fields that are more similar to the target field i in terms of crop climate sensitivity. .

[0030] Based on this, for target field i, calculate the neighboring fields. The initial interpolated climate deviation relative to target field i during this historical growing season. Taking temperature as an example, the temperature deviation... ,in For adjacent fields The initial interpolated temperature sequence, The initial interpolated temperature sequence for target field i; precipitation deviation. , sunshine deviation Calculate in the same way. Utilize spatial influence weights. Temperature deviation We obtain the temperature correction factor for target field i by weighted summation: In the formula, For spatial influence weights, Temperature deviation, precipitation correction factor and sunlight correction factor Calculate using the exact same method, and separately calculate the temperature deviation in the formula. Replace with precipitation deviation and solar radiation deviation That is, all variables in the formula have the same meaning as above. The correction factor is directly superimposed onto the initial interpolated climate of target field i to obtain the temperature, precipitation, and sunshine components in the historical effective climate driving factor sequence: In the formula, , and These are the effective temperature sequence, effective precipitation sequence, and effective sunshine sequence for target field i during the historical growing season, respectively, with units consistent with the initial sequence. The initial interpolated temperature sequence for target field i; Temperature correction factor; The initial interpolated precipitation sequence for target field i; Precipitation correction factor; The initial interpolated sunshine sequence for target field i; This is the sunshine correction factor. After the above correction, the historical effective climate driving factor sequence integrates data from neighboring fields. The spatial deviation information was obtained, and the deviation contribution was directionally filtered and weighted through the climate response feature vector, so as to more closely reflect the actual climate stress situation felt by the crop in the field during the historical growing season.

[0031] S3: Use the statistics of climate driving factor sequences, historical soil available nutrients, and yield and quality data as the training set to train a Bayesian multi-output response surface model.

[0032] After completing the spatial correction of the interpolated climate sequences for each field and each historical growing season and generating historical effective climate driving factor sequences, the current need is to establish a predictive model that can describe the quantitative relationship between yield, tobacco quality, and climate, soil, and fertilizer application. The response of tobacco yield and quality in hilly and mountainous tobacco-growing areas to nitrogen, phosphorus, and potassium fertilizers exhibits significant nonlinear characteristics, and there are also complex interaction effects between climate factors and basic soil fertility. Traditional univariate or multivariate quadratic fertilizer effect equations are insufficient to simultaneously characterize the nonlinear mapping of multiple input variables to multiple quality indicators and the drift of this mapping with changes in climate conditions. Therefore, based on historical effective climate driving factor sequences, historical soil available nutrients, and yield and quality data, a Bayesian multi-output response surface model is trained to capture this nonlinear response relationship and retain predictive uncertainty using probabilistic modeling.

[0033] When constructing the training set, statistics were first extracted for each historical effective growing season of each field within the region. Based on the growth process characteristics of tobacco cultivation, the effective temperature sequence of a single growing season was then analyzed. Effective precipitation sequence and effective sunshine sequence The growing season is divided into three consecutive growth stages: the rooting period (days 1-30 after transplanting), the vigorous growth period (days 31-70), and the maturity period (days 71-the end of the harvest season). The number of days for each stage can be adjusted by the implementer according to the specific developmental rhythm of the local variety, but once set, it must be used uniformly throughout all historical seasons. For each growth stage, three statistics are calculated: the average daily effective temperature within the stage, in degrees Celsius (°C); the cumulative daily effective precipitation within the stage, in millimeters (mm); and the average daily effective sunshine hours within the stage, in hours (h). Thus, a total of nine climate statistics are obtained for one growing season, forming a nine-dimensional climate-driven statistical vector. .

[0034] Soil available nutrient data measured before sowing in the same historical growing season were included as input variables. Soil available nitrogen content was collected. The unit is milligrams per kilogram (mg / kg), representing the available phosphorus content in the soil. The unit is milligrams per kilogram (mg / kg), indicating the available potassium content in the soil. The unit is milligrams per kilogram (mg / kg). Simultaneously, the actual pure nitrogen fertilizer applied during the season should be recorded. phosphate fertilizer purity and potassium fertilizer purity The units are all kilograms per hectare (kg·hm). -2 The climate-driven statistical vector, three soil nutrient contents, and three actual applied pure amounts are concatenated into a 15-dimensional input feature vector. The output target for this season is the measured yield per unit area of ​​tobacco leaves after harvest in this field. (Unit: kg·hm) -2 Nicotine content (Percentage by mass, dimensionless), Total sugar content (mass percentage, dimensionless) and potassium content (mass percentage, dimensionless), forming the output vector Samples from all fields within the region across all valid historical growing seasons were compared. The data is combined to form a training set. ,in This represents the total number of valid historical samples.

[0035] The dimensions of the input feature vectors vary considerably, and directly inputting them into a Gaussian process model can cause certain dimensions to dominate the covariance structure during kernel function calculation. Therefore, each dimension of all input feature vectors in the training set is standardized: the mean of that dimension is calculated over all samples. and standard deviation Then, perform a dimension-by-dimensional transformation on each sample. The standardized input matrix is ​​denoted as... The corresponding output matrix is ​​still denoted as The mean and standard deviation parameters used are preserved so that the exact same linear transformation is performed on new inputs when the model makes predictions.

[0036] This embodiment employs Bayesian Gaussian process regression to construct a multi-output response surface model. Gaussian process regression is a well-known technique that defines the distribution of functions in a function space in the form of probabilistic priors. It can output the predicted mean and predicted variance, directly characterizing the uncertainty of the prediction. Considering that yield and the three quality indicators have independent observation noise under given environmental and management conditions, and that the interaction between the quality indicators in actual production has been introduced through input variables, to simplify the model structure and computational cost, the yield per unit area of ​​tobacco leaves is used as the model variable. Nicotine content Total sugar content and potassium content Each of the four sub-models trains an independent Gaussian process regression model, and each sub-model shares the standardized input feature vector. However, its kernel function parameters and noise variance are not shared and are learned independently.

[0037] Taking the output sub-model as an example, its modeling form is as follows: ,in Given a zero-mean Gaussian process prior, It follows the principle of zero mean and variance. Gaussian white noise. The prior covariance is defined by the squared exponential kernel function: In the formula, and The input feature vectors are standardized from the two training samples; Let V be the signal variance, and let A be the overall fluctuation amplitude of the control function. This is a length scale parameter that controls the smoothness of function value changes caused by input variations; The kernel function represents the Euclidean norm. This kernel function indicates that the smaller the Euclidean distance between two input points in the feature space, the stronger the correlation between their corresponding function values. This aligns with the physical prior that adjacent input points in a continuous agricultural production system have similar yield responses. Nicotine content. Total sugar content and potassium content The sub-models use the exact same kernel function form and independently hold their own signal variances. Length scale parameters and noise variance Preset parameter set.

[0038] The training process for each sub-model is achieved by maximizing the log-marginal likelihood of its output variable on the training set. Learning the log-marginal likelihood automatically balances goodness of fit with model complexity, without relying on manually set penalty coefficients. The optimization algorithm employs either the conjugate gradient method or L-BFGS (Limited-memory Broyden–Fletcher–Goldfarb–Shanno), performing multiple optimizations starting from several sets of random initial values, and selecting the preset parameter combination that maximizes the log-marginal likelihood as the training result. After training, each sub-model retains the observed training inputs. and its corresponding output vector And the optimized kernel function preset parameters and noise variance For any new standardized input vector The conditional prediction distribution of the sub-model is a Gaussian distribution, and its prediction mean and prediction variance are given by the well-known Gaussian process prediction formula. The prediction mean serves as the expected estimate of the yield or quality index under the input conditions, and the prediction variance reflects the degree of uncertainty of the estimate.

[0039] When facing deployment environments with significantly limited computing resources, the Bayesian Gaussian process regression in this step can be reduced to Bayesian multinomial regression. In this case, a fixed-order multinomial basis function is used to explicitly expand the input features, and Bayesian priors are assigned to the expansion coefficients. The predicted distribution is then obtained through posterior inference. Bayesian multinomial regression can also provide probabilistic predictions and has lower computational cost when the feature dimension is low. However, its adaptability to complex nonlinearities is weaker than that of Gaussian process regression, making it a suitable alternative.

[0040] Through the above training, a system was obtained with climate-driven statistics, available soil nutrients, and fertilizer application rate as inputs, and tobacco yield per unit area as input. Nicotine content Total sugar content and potassium content This is a Bayesian multi-output response surface model that provides predicted means and variances as outputs. This model will be stored as a foundational probabilistic predictor for fertilization optimization decisions in subsequent production seasons within the region. When the inputs are based on the current production season's effective climate driving factor statistics, measured soil nutrients, and candidate fertilization schemes, the model can be directly invoked to obtain the predicted distributions of each output indicator.

[0041] S4: In the current production season, generate the effective climate driving factor sequence for the current field in the same way, extract its statistics, and input them together with the current available soil nutrients and the set candidate fertilizer amount into the Bayesian multi-output response surface model to jointly predict yield and multi-dimensional quality indicators. Calculate the quality coordination index based on the preset quality target and weight, and generate the optimal fertilization plan using a multi-objective evolutionary algorithm and the ideal point method.

[0042] For target field i, firstly, using real-time data from meteorological observation stations and short-term numerical weather prediction products, conventional kriging interpolation is used to obtain the daily initial temperature, initial precipitation, and initial sunshine estimation sequences for the current growing season. The most recently maintained climate response feature vector for the field is retrieved from the field archive; if the field has not yet accumulated a valid feature vector, the regional average vector is used. The target field is then used as the center, with a radius of... Determine the set of neighboring fields within the range. An empirical value of 3km is used as a preset parameter, but it can be adjusted by the implementer according to the specific implementation scenario. Following the same deviation correction procedure as the generation of the historical effective climate driving factor sequence, calculations are performed for each neighboring field. The cosine similarity of the target field i to the climate response feature vector is non-negatively truncated and then sharpened by a sharpening factor. After exponentiation and normalization, the spatial influence weights are obtained. , The empirical value of 2 is taken; it is dimensionless and serves as a preset parameter, but it can also be adjusted. This space is used to influence the weights. For adjacent fields The temperature correction factor is obtained by weighted summation of the initial interpolated climate bias relative to target field i. Precipitation correction factor and sunlight correction factor The correction factor is then superimposed onto the initial interpolation sequence to generate the effective temperature sequence for the current field during the current growing season. Effective precipitation sequence and effective sunshine sequence After obtaining the effective climate driving factor sequence, the effective temperature sequence was divided into fertility stage segments using the same method as the training stage. Effective precipitation sequence and effective sunshine sequence The period was divided into three stages: root extension (days 1-30 after transplanting), vigorous growth (days 31-70), and maturity (days 71 to the end of harvest). For each stage, the daily average temperature (°C), cumulative precipitation (mm), and daily average sunshine duration (h) were calculated, resulting in a total of nine statistics, which were combined into a nine-dimensional climate-driven statistical vector. Simultaneously, soil samples were collected from the target fields before transplanting in the current season, and the available nitrogen content in the soil was determined by routine laboratory analysis. (Unit: mg / kg), available phosphorus content (Unit: mg / kg) and available potassium content (Unit: mg / kg). Define the continuous feasible region for the fertilization decision variable for the current season: pure nitrogen application rate. The range of values ​​is upper limit Get experience points ; Pure phosphorus application The range of values ​​is upper limit Get experience points Potassium application rate The range of values ​​is upper limit Get experience points The upper bounds are preset parameters, which can also be adjusted by the implementer according to the specific implementation scenario. The target values ​​and weighting coefficients for each quality indicator are preset by the tobacco purchasing company: Target value for nicotine content. Take an empirical value of 2.0% (by mass) as the target value for total sugar content. Take the empirical value of 20.0% (mass percentage) as the target value for potassium content. Take an empirical value of 1.5% (quality percentage); the corresponding quality target weighting coefficient. , , All values ​​are taken as empirical values ​​of 1.0, dimensionless. The target values ​​and weighting coefficients mentioned above are preset parameters, but can be adjusted by the implementer according to the specific implementation scenario. For any set of fertilizer application rates... Compare it with the current climate-driven statistical vector and readily available nutrients in the soil , , The input feature vectors are merged into a 15-dimensional vector, and the mean values ​​of each dimension saved during the training phase are utilized. and standard deviation Standardize to obtain a standardized input vector. Standardize the input vector Input the four pre-trained independent Gaussian process regression sub-models respectively to obtain the mean yield prediction under the current input conditions. (Unit: kg·hm) -2 ), predicted average nicotine content (Percentage by mass), predicted mean of total sugar content (Percentage by mass) and predicted mean of potassium content (Quality percentage). Based on the predicted quality indicators, calculate the quality consistency index: In the formula: , , These are the predicted values ​​for nicotine content, total sugar content, and potassium content, respectively. , , These are the pre-set quality target values ​​for the corresponding enterprises; , , These are the weighting coefficients for each quality objective; Divide zero by the protection constant, and (For example The higher the value of this indicator, the smaller the overall deviation between the predicted quality and the target quality, and the better the overall coordination. (Based on nitrogen application rate) Phosphorus application rate Potassium application rate For the decision variables, within the aforementioned feasible region, the Non-Dominated Sort Genetic Algorithm (NSGA-II) is used for multi-objective evolutionary solving. NSGA-II is a well-known multi-objective evolutionary algorithm. In this embodiment, its population size is set to 100, the number of generations to 200, the simulated binary crossover probability to 0.9, and the polynomial mutation probability to 0.1, all of which are preset parameters and can be adjusted. During the optimization process, each individual represents a set of fertilization schemes. Its fitness evaluation process is as follows: the fertilization amount corresponding to the individual is combined with the current climate-driven statistics and soil available nutrients in the aforementioned manner and standardized, then input into the Bayesian response surface model to obtain the predicted yield. And the quality forecast mean, and then calculate the quality consistency. The individual's values ​​for the two objectives are... The algorithm automatically performs non-dominated sorting and crowding calculation, and after multiple generations of evolution, outputs the Pareto front solution set. Each solution is a non-dominated fertilization scheme vector and its corresponding predicted yield and quality coordination values. Obtain the Pareto front solution set. Next, selection was based on fertilizer cost constraints. A maximum limit for the total net fertilizer application rate throughout the entire growth period was preset. Take the empirical value of 300 kg·hm -2 These are preset parameters, but can also be adjusted by the implementer according to the specific implementation scenario. For the Pareto front solution set... For each solution, calculate the sum of its three elements: scalars. ,like Greater than Then from the Pareto front solution set By removing this solution, we obtain a set of feasible solutions that satisfy the cost constraints. In the feasible solution set In order to predict output and quality coordination As a two-dimensional evaluation coordinate, determine the positive ideal point. ,in , Calculate the feasible solution set. Normalized Euclidean distance from each solution to the positive ideal point: In the formula, For quality consistency, For the first The predicted output corresponding to the group scheme, This represents the ideal value for the production target. Let be the quality coordination and scheduling index corresponding to the j-th scheme. The ideal value for quality coordination and scheduling objectives. Divide by zero by the protection constant, and take the value of The deviations are normalized using the positive ideal values ​​of each target to eliminate differences in the dimensions and numerical ranges of the output and coordination indicators. The first [target] is selected. Normalized Euclidean distance from each solution to the ideal point The solution with the smallest overall distance is taken as the optimal compromise, and the corresponding nitrogen, phosphorus, and potassium application rates are output as the optimal fertilization plan for the current production season. The ideal point method is a well-known multi-attribute decision-making method and will not be elaborated on here.

[0043] S5: After each growing season, collect measured yield and quality data, remove outliers after validity testing, add the measured yield and quality data as new samples to the training set, update the parameters of the Bayesian multi-output response surface model by re-maximizing the log marginal likelihood, and smoothly update the currently used climate response feature vector based on the similarity between the old and new vectors when the sample size condition is met. The update results are used for decision-making in the next season.

[0044] After the current production season of the target field ends, retrieve the actual net amounts of nitrogen, phosphorus, and potassium fertilizers applied during the season from the field records, and collect field measurement data: measured yield per unit area of ​​tobacco leaves. (Unit: kg·hm) -2 ), measured nicotine content (mass percentage), measured total sugar content (percentage by mass) and measured potassium content (Percentage of quality). Simultaneously, retrieve the statistics of effective climate driving factors generated this season. Compared with the soil available nutrient content measured before transplanting , , The above data, together with the measured yield and quality indicators, constitute the feedback sample for this season.

[0045] Before feedback samples are used for updates, data validity checks must be performed because abnormal field conditions (such as severe pests and diseases, extreme weather disasters, or sampling errors) may cause measured values ​​to deviate from the normal range. For the yield dimension, check the number of valid historical yield records accumulated in the target field: if it has at least three growing seasons of historical yield records, then read its lowest historical yield. and highest historical production When the actual yield satisfy or When the feedback sample is deemed abnormal, the coefficients 0.5 and 1.5 are used as the first and second preset percentages, respectively, taking empirical values ​​of 50% and 150%, dimensionless, and are preset parameters that can be adjusted by the implementer according to the specific implementation scenario. If the historical yield record of the field is less than three growing seasons, the historical minimum and historical maximum yields of all fields in the planting area where the field is located are used as reference boundaries to apply the same percentage threshold judgment. For the quality dimension, the measured nicotine content, total sugar content, and potassium content are checked to see if they exceed the preset physiologically feasible limits for tobacco leaves: nicotine content below 0.5% or above 5.0%, total sugar content below 5% or above 35%, and potassium content below 1% or above 8%. These limits are preset parameters and can be adjusted. When any quality indicator exceeds the corresponding limit, the feedback sample is also deemed abnormal. Samples deemed abnormal do not participate in any subsequent update calculations and are only recorded in the abnormal event log for auditing; only samples that pass both yield and quality tests are marked as valid samples, constituting the valid new evidence set for this season.

[0046] For parameter updates in the Bayesian multi-output response surface model, since each sub-model is constructed using Bayesian Gaussian process regression, the posterior distribution of its preset parameters can naturally evolve according to Bayesian rules after receiving new evidence. Specifically, the input-output data of the effective samples in this quarter are compared... Add the historical training set of the corresponding sub-model, where Climate-driven statistics for this season The available nutrients in the soil and the actual amount of fertilizer applied were obtained after standardization using the preserved mean and standard deviation. Actual production Actual nicotine content Measured total sugar content and measured potassium content The component corresponding to this sub-model. For the measured yield. Actual nicotine content Measured total sugar content and measured potassium content The four sub-models update the signal variance of each model by either re-maximizing the log-marginal likelihood based on the expanded training set or approximating the posterior mean of the preset parameters using lightweight variational inference. Length scale and noise variance In this process, historical prior information and new likelihood are automatically balanced in the optimization objective. The influence of single-season samples on the preset parameters is constrained by the amount of existing data, avoiding drastic shifts in the model's predicted distribution. The updated preset parameters and the expanded training set are persistently saved as the basis for the Bayesian response surface model when optimizing fertilization in the next season.

[0047] The update of the climate response feature vector is also based solely on the growing season data corresponding to the valid samples of this season. This involves the daily valid temperature series of the valid growing season. Effective precipitation sequence and effective sunshine sequence and measured yield Actual nicotine content Measured total sugar content Add to the historical dataset of the target field to reassess the cumulative number of effective growing seasons for that field since the start of recording. When the effective growing season quantity Less than the preset minimum sample size threshold hour, An empirical value of 5 is used as the preset parameter, but it can be adjusted by the implementer according to the specific implementation scenario. Because re-performing partial least squares regression under insufficient sample size will introduce high variance in the estimate, the current climate response feature vector is maintained in this case. If unchanged, no subsequent vector update process will be executed. At that time, based on the entirety of the field Using data from each effective growing season, and following the same partial least squares regression procedure as S1, a new climate response feature vector was extracted. Because the tobacco growing environment is a non-stationary system, interannual climate fluctuations and soil variability may cause differences in the directions of old and new climate response vectors. To measure the degree of this difference, a new climate response characteristic vector is calculated. With the current climate response feature vector Cosine similarity between them: In the formula, For vector dot product, For the Euclidean norm, The range of values ​​is , dimensionless, the smaller its value, the greater the difference in climate response patterns represented by the old and new vectors.

[0048] Set mutation threshold The empirical value of 0.7 is used; it is dimensionless and is a preset parameter, but it can also be adjusted. If... This indicates that the seasonal data induced a large directional perturbation in the climate response characteristics. To suppress abrupt changes, the smoothing coefficient was adjusted. Reduced from the default value to The default smoothing coefficient is an empirical value of 0.3, and after reduction, it is an empirical value of 0.1. These are preset parameters and can be adjusted. The smoothing coefficient remains at its default value. Then, an exponential moving average is used to smoothly update the climate response feature vector: In the formula, This is the updated climate response feature vector. This is the new climate response feature vector. This is the current climate response feature vector. The smoothing coefficient is used. This vector is smoothed... After norm normalization, the vectors are written into the target field's file, replacing the original vectors, and serve as the basis for generating effective climate driving factor sequences during spatial correction in the next production season. Following the above validity checks, Bayesian posterior updates, and protected conditional vector smoothing updates, the model parameters and climate response feature vectors absorb effective information from the current season in a controlled manner. The updated results are directly used in the fertilization decision-making process for the next season.

[0049] This application also discloses a tobacco production decision-making system that integrates climate data and fertilization strategies, including a processor and a memory. The memory stores computer program instructions, which, when executed by the processor, implement the tobacco production decision-making method that integrates climate data and fertilization strategies according to this application.

[0050] The system also includes other components well known to those skilled in the art, such as communication buses and communication interfaces, the settings and functions of which are known in the art and will not be described in detail here.

[0051] In this application, the aforementioned memory can be any tangible medium that contains or stores a program that can be used or combined with an instruction execution system, apparatus, or device. For example, a computer-readable storage medium can be any suitable magnetic or magneto-optical storage medium, such as resistive random access memory, dynamic random access memory, static random access memory, etc., or any other medium that can be used to store required information and can be accessed by an application program, module, or both.

[0052] The above are merely preferred embodiments of the present invention. It should be noted that those skilled in the art can make several improvements and substitutions without departing from the technical principles of the present invention, and these improvements and substitutions should also be considered within the scope of protection of the present invention.

Claims

1. A tobacco production decision-making method integrating climate data and fertilization strategies, characterized in that, include: Historical interpolated climate sequences and yield and quality data for each field are obtained. Partial least squares regression is used to extract the climate response feature vector of the field. Based on the climate response feature vector, the cosine similarity between the target field and neighboring fields is calculated to construct spatial influence weights. The spatial influence weights are used to correct the deviation of the historical interpolated climate sequences to generate a historical effective climate driving factor sequence. The statistics of the climate driving factor sequence, historical soil available nutrients, and yield and quality data are used as the training set to train a Bayesian multi-output response surface model. In the current production season, the effective climate driving factor sequence for the current field is generated in the same manner, its statistics are extracted, and input into the Bayesian multi-output response surface model along with the current soil available nutrients and the set candidate fertilization amounts. This jointly predicts yield and multidimensional quality indicators. Based on preset quality targets and weights, the quality coordination index is calculated, and the optimal fertilization scheme is generated using a multi-objective evolutionary algorithm and the ideal point method. After each growing season, measured yield and quality data are collected, and outliers are removed after validity testing. The measured yield and quality data are added to the training set as new samples. The parameters of the Bayesian multi-output response surface model are updated by re-maximizing the log-marginal likelihood. The currently used climate response feature vector is smoothly updated based on the similarity between the old and new vectors when the sample size condition is met. The updated results are used for decision-making in the next season.

2. The tobacco production decision-making method integrating climate data and fertilization strategies according to claim 1, characterized in that, The historical interpolated climate series includes daily temperature estimation series, daily precipitation estimation series, and daily sunshine duration estimation series; the yield is the yield per unit area of ​​tobacco leaves, and the quality data includes nicotine content and total sugar content; the extraction of the climate response feature vector of the field using partial least squares regression includes: using the daily temperature estimation series, daily precipitation estimation series, and daily sunshine duration estimation series as independent variables, and the yield per unit area of ​​tobacco leaves, the nicotine content, and the total sugar content as multiple response dependent variables, extracting a standardized regression coefficient vector as the climate response feature vector; when the historical interpolated climate series of a certain field and the yield and quality data are less than a preset number of growing seasons, the average climate response feature vector of all fields in the region is used as the climate response feature vector of that field.

3. The tobacco production decision-making method integrating climate data and fertilization strategies according to claim 1, characterized in that, The construction of spatial influence weights includes: taking the target field as the center, forming a set of neighboring fields within a preset radius; calculating the cosine similarity between the target field and each neighboring field in the set on the climate response feature vector; truncating the cosine similarity non-negatively, and normalizing the non-negatively truncated cosine similarity by exponentiation with a preset sharpening factor to obtain the normalized spatial influence weights.

4. The tobacco production decision-making method integrating climate data and fertilization strategies according to claim 3, characterized in that, The step of using the spatial influence weights to correct the deviation of historical interpolated climate and generate a historical effective climate driving factor sequence includes: calculating the initial interpolated climate deviation of each neighboring field relative to the target field; using the spatial influence weights to perform a weighted summation of the initial interpolated climate deviation to obtain a correction factor; and superimposing the correction factor onto the initial interpolated climate of the target field to obtain the historical effective climate driving factor sequence; wherein, the historical effective climate driving factor sequence includes an effective temperature sequence, an effective precipitation sequence, and an effective sunshine sequence.

5. The tobacco production decision-making method integrating climate data and fertilization strategies according to claim 1, characterized in that, The statistical measures for extracting the climate driving factor sequence include: dividing the historical effective climate driving factor sequence into the rooting period, the vigorous growth period, and the maturity period according to the tobacco growth period; extracting the average temperature, cumulative precipitation, and average sunshine duration for each period, and combining them to form the statistical measures.

6. The tobacco production decision-making method integrating climate data and fertilization strategies according to claim 1, characterized in that, The Bayesian multi-output response surface model is either a Bayesian Gaussian process regression model or a Bayesian multinomial regression model. The input variables of the Bayesian multi-output response surface model include the statistics, soil available nitrogen content, soil available phosphorus content, soil available potassium content, nitrogen application rate, phosphorus application rate, and potassium application rate. The output variables include tobacco yield per unit area, nicotine content, total sugar content, and potassium content.

7. The tobacco production decision-making method integrating climate data and fertilization strategies according to claim 1, characterized in that, The quality coordination index is defined as follows: the square of the difference between the predicted value of each quality index in the multidimensional quality index and its corresponding enterprise preset quality target value is calculated; the square deviations are weighted and summed using the enterprise preset quality target weight coefficients; the weighted sum is added to a zero division protection constant and the reciprocal is taken to obtain the quality coordination index, such that the larger the value of the quality coordination index, the better the overall quality coordination.

8. The tobacco production decision-making method integrating climate data and fertilization strategies according to claim 7, characterized in that, The method of generating the optimal fertilization scheme using a multi-objective evolutionary algorithm and the ideal point method includes: using nitrogen, phosphorus, and potassium application rates as decision variables, within a preset fertilization rate constraint range, employing a fast non-dominated sorting genetic algorithm to simultaneously maximize the predicted yield and the quality coordination index, thereby obtaining a Pareto front solution set; and combining fertilizer cost constraints, using the ideal point method to select the compromise fertilization scheme with the optimal comprehensive distance from the Pareto front solution set, as the optimal fertilization scheme.

9. The tobacco production decision-making method integrating climate data and fertilization strategies according to claim 1, characterized in that, The validity check includes: when the target field has a historical yield record of no less than a preset number of growing seasons, if the measured yield is lower than a first preset percentage of its lowest historical yield or higher than a second preset percentage of its highest historical yield, then the sample corresponding to the measured yield is determined to be abnormal data; when the target field's historical yield record does not meet the preset historical data accumulation requirements, the lowest and highest historical yields of all fields in the planting area where the field is located are used as reference boundaries for determination; if any item in the measured quality data exceeds the preset physiologically feasible limit for tobacco leaves, the corresponding sample is also determined to be abnormal data; data determined to be abnormal is removed and does not participate in subsequent update calculations.

10. A tobacco production decision-making system integrating climate data and fertilization strategies, characterized in that, include: Processor and memory; The memory stores computer program instructions that, when executed by the processor, implement the tobacco production decision-making method according to any one of claims 1 to 9, which integrates climate data and fertilization strategies.