A method and system for predicting crop yield based on multi-data fusion
By using a multi-data fusion approach, configuring multi-level farmland monitoring standards, and utilizing an intelligent control system and ARIMAX model, the problem of low accuracy in crop yield prediction in existing technologies has been solved, achieving higher precision yield prediction.
Patent Information
- Application Number
- CN202411782338.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-05
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-12-05
AI Technical Summary
Current crop yield prediction technologies rely on a single data source and simple empirical models, ignoring the complex factors and their interactions during crop growth, resulting in low accuracy and reliability of prediction results.
A multi-data fusion method is adopted, multi-level farmland monitoring standards are configured, and data from multiple monitoring devices are acquired using an intelligent control system. Feature extraction and dimensionality reduction are performed using the PCA algorithm, deviation values are calculated, and an ARIMAX model is constructed to predict crop yield.
It improves the accuracy and reliability of crop yield forecasting, and can more accurately consider the response of crops to environmental changes at different growth stages, providing more accurate yield forecast data.
Smart Images

Figure CN119558484B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of agricultural information processing technology, and in particular to a method and system for predicting crop yield based on multi-data fusion. Background Technology
[0002] In modern agriculture, accurate crop yield forecasting is of paramount importance. With the continuous growth of the global population, the demand for agricultural products is constantly increasing, requiring more scientific and efficient management methods for agricultural production. Crop yield forecasting helps farmers, agricultural enterprises, and relevant government departments to plan resource allocation in advance, formulate reasonable market strategies, and ensure a stable supply of agricultural products. Traditional crop yield forecasting methods often rely on a single data source or simple empirical models, ignoring the numerous complex factors and their interactions during crop growth. They fail to accurately consider the different responses of crops to environmental changes at different growth stages and lack the ability to comprehensively consider factors such as soil quality and pesticide residues. Therefore, their accuracy and reliability are low, and they cannot provide accurate yield forecast data for enterprises. These models have relatively simple structures, consider fewer factors, and their accuracy needs improvement. Traditional crop yield estimation methods mainly include agronomic forecasting methods, statistical forecasting methods, and meteorological forecasting methods. Summary of the Invention
[0003] This application provides a method and system for predicting crop yield based on multi-data fusion, aiming to solve the technical problem of low accuracy and reliability in existing technologies that rely on single data sources and simple empirical models. To further improve crop yield prediction performance, the method of this invention is easy to implement and highly effective.
[0004] According to a first aspect of this disclosure, a method for predicting crop yield based on multi-data fusion is provided, including:
[0005] For the defined multi-level growth stages of crops, multi-level farmland monitoring standards are configured and integrated into an intelligent control system. These multi-level growth stages include germination, seedling, growth, and maturity. Based on these stages, multi-source monitoring equipment data is periodically acquired, with the equipment communicating with the intelligent control system. The acquired multi-source monitoring equipment data is processed using PCA algorithm for feature extraction and dimensionality reduction, resulting in dimensionality-reduced data. This data is then used to update the crop growth element set, which is constructed based on the content of the multi-level farmland monitoring standards and includes soil quality data related to crop growth. The system comprises a set of quantitative elements, a set of weather conditions, and a set of pesticide residues. A sliding window is set, with the sliding step size determined according to the multi-level growth stages of the crop. The average value of the data from each monitoring device updated within each element set of the crop growth element set within the window is calculated. The average value of the data from each monitoring device within each element feature set is compared with the corresponding monitoring standard in the multi-level farmland monitoring standard to calculate the deviation value, thus obtaining a multi-source deviation dataset. Based on the sliding step size, the multi-source deviation data within each window are arranged sequentially according to the sliding step length to form a time-series multi-source deviation dataset. The time-series multi-source deviation dataset is input into a crop growth prediction model to obtain a predicted crop yield value. The prediction model is constructed using an ARIMAX architecture.
[0006] According to a second aspect of this disclosure, a crop yield prediction system based on multi-data fusion is provided, comprising:
[0007] The system includes a multi-level farmland monitoring standard configuration module, which configures multi-level farmland monitoring standards for defined multi-level crop growth stages and integrates these standards into the intelligent control system. The multi-level growth stages include germination, seedling, growth, and maturity. A multi-source monitoring equipment data acquisition module is used to periodically acquire multi-source monitoring equipment data about the farmland environment based on the defined crop growth stages. The multi-source monitoring equipment is communicatively connected to the intelligent control system. A multi-source monitoring equipment data feature extraction module uses the PCA algorithm to extract features and reduce the dimensionality of the acquired multi-source monitoring equipment data, obtaining dimensionality-reduced multi-source monitoring equipment data and updating the crop growth element set. This crop growth element set is constructed based on the content of the multi-level farmland monitoring standards and includes soil conditions related to crop growth. The system comprises: a soil quality element set, a weather condition element set, and a pesticide residue element set; a multi-source deviation dataset acquisition module, which sets a sliding window with a sliding step size according to the multi-level growth stages of the crop, calculates the average value of each monitoring device data updated in each element set of the crop growth element set within the window, compares the average value of each monitoring device data in each element feature set with the corresponding monitoring standard in the multi-level farmland monitoring standard, calculates the deviation value, and obtains the multi-source deviation dataset; a time-series multi-source deviation dataset construction module, which arranges the multi-source deviation data in each window according to the sliding step size to form a time-series multi-source deviation dataset; and a crop yield prediction value acquisition module, which inputs the time-series multi-source deviation dataset into a crop growth prediction model to obtain the crop yield prediction value, wherein the prediction model is built using ARIMAX architecture.
[0008] One or more technical solutions provided in this disclosure have at least the following technical effects or advantages: For the defined multi-level growth stages of crops, multi-level farmland monitoring standards are configured, and these standards are configured in an intelligent control system. The multi-level growth stages include germination, seedling, growth, and maturity. Based on these multi-level growth stages, data from multi-source monitoring devices monitoring the farmland environment are periodically acquired. These multi-source monitoring devices are communicatively connected to the intelligent control system. The acquired multi-source monitoring device data is processed using a PCA algorithm for feature extraction and dimensionality reduction to obtain dimensionality-reduced multi-source monitoring device data. The crop growth element set is then updated based on the content of the multi-level farmland monitoring standards. The growth element set includes soil quality elements, weather condition elements, and pesticide residue elements related to crop growth. A sliding window is set, with the sliding step size set according to the multi-level growth stages of the crop. The average value of the data updated by each monitoring device in each element set of the crop growth element set within the window is calculated. The average value of the data from each monitoring device in each element feature set is compared with the corresponding monitoring standard in the multi-level farmland monitoring standard to calculate the deviation value and obtain a multi-source deviation dataset. Based on the sliding step size, the multi-source deviation data in each window are arranged in chronological order according to the sliding step length to form a time-series multi-source deviation dataset. The time-series multi-source deviation dataset is input into a crop growth prediction model to obtain the predicted crop yield value, wherein the prediction model is built using ARIMAX architecture. This solves the technical problem of existing technologies that rely on a single data source or simple empirical models, ignoring the numerous complex factors and their interactions in the crop growth process, and achieves the technical effect of improving prediction accuracy.
[0009] The above description is merely an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, specific embodiments of this application are given below. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0011] Figure 1 A flowchart illustrating the crop yield prediction method based on multi-data fusion provided in this application embodiment;
[0012] Figure 2A schematic diagram of the structure of a crop yield prediction system based on multi-data fusion provided in an embodiment of this application.
[0013] Figure labeling: Module 11 for standard configuration of multi-level farmland monitoring, Module 12 for data acquisition of multi-source monitoring equipment, Module 13 for feature extraction of data from multi-source monitoring equipment, Module 14 for obtaining multi-source deviation dataset, Module 15 for constructing time-series multi-source deviation dataset, and Module 16 for obtaining predicted crop yield. Detailed Implementation
[0014] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0015] Example 1
[0016] The crop yield prediction method based on multi-data fusion provided in this disclosure is referred to below. Figure 1 The methods include:
[0017] S100: For the defined multi-level growth stages of crops, configure multi-level farmland monitoring standards and configure the multi-level farmland monitoring standards in the intelligent control system, wherein the multi-level growth stages include germination period, seedling period, growth period and maturity period.
[0018] Furthermore, the multi-level farmland monitoring standard includes: configuring a first monitoring standard, which is based on soil quality at multiple growth stages of crops;
[0019] Configure a second monitoring standard, which is based on weather conditions at multiple growth stages of crops;
[0020] Configure a third monitoring standard, which is based on pesticide residues at multiple growth stages of crops;
[0021] By performing an index interaction correlation analysis on the first monitoring standard, the second monitoring standard, and the third monitoring standard, multi-level farmland monitoring standards are determined.
[0022] Specifically, based on the comprehensive results of plant physiology, agricultural science, and crop cultivation practices, plant growth stages are divided into multiple levels, including germination, seedling, growth, and maturity. Multi-level farmland monitoring standards for various crops at each stage are defined based on historical data from long-term observations in the region, research findings of agricultural scientists, experience summaries from agricultural extension stations and agricultural technicians, and international and domestic agricultural standards. These standards include soil quality, weather conditions, and pesticide residues. Soil quality is reflected by soil pH, and weather conditions are reflected by temperature, humidity, and light intensity. The specific multi-level farmland monitoring standard values for each growth stage are as follows.
[0023] Multi-level farmland monitoring standards during germination include: a suitable soil pH of 6.0-7.0, which is conducive to the activity of enzymes in the seeds and promotes seed germination; a suitable temperature of 12-20°C, which helps seeds break dormancy and begin the germination process; a suitable humidity of 50%-60%, as too little water makes germination difficult and too much water may lead to oxygen deficiency and seed rot; for crops that require light, a light intensity of 1000-1500 lux is sufficient, because seeds mainly rely on their own stored nutrients in the early stages of germination, and light is not a key factor; excessive light may even inhibit germination; and pesticide residues should be as low as possible, as seeds and seedlings are very sensitive to pesticides during germination, and even trace amounts of pesticide residues can seriously affect their growth and development. Therefore, before sowing, it should be ensured that there are no pesticide residues in the soil that may affect seed germination.
[0024] Multi-level farmland monitoring standards during the seedling stage include: a suitable soil pH of 6.0-7.5, which facilitates nutrient absorption by seedling roots; a suitable temperature of 15-22°C, which promotes root growth and above-ground development; a humidity range of 55%-65%, which meets the seedling's water requirements while ensuring good soil aeration for root respiration; a light intensity range of 1500-3000 lux, as seedlings require adequate light for photosynthesis; and pesticide residue standards that are as low as possible and meet strict safety standards. Seedlings have relatively weak tolerance to pesticides, and high concentrations of pesticide residues should be avoided to prevent adverse effects on their growth and development.
[0025] Multi-level farmland monitoring standards during the growing season include: a suitable soil pH of 6.5-7.5, which promotes the growth of rhizobia and nitrogen fixation, and also facilitates the absorption of other nutrients by crops; a temperature range of 18-25°C, which is conducive to the metabolism of crops during their vigorous growth period, including nutrient absorption and photosynthesis; a humidity range of 60%-70%, ensuring sufficient soil moisture for crop growth, but avoiding waterlogging to maintain good soil physical structure and aeration; a light intensity range of 2000-5000 lux, as the demand for light intensity increases with crop growth, requiring the energy needs of photosynthesis to be met and the accumulation of photosynthetic products to be increased; and pesticide residues must comply with agricultural product safety production standards. During the growing season, pest and disease control measures may be implemented, necessitating strict control of pesticide usage and safe intervals to ensure that pesticide residues are within safe limits at harvest.
[0026] Multi-level farmland monitoring standards during the ripening period include: soil pH between 6.0 and 7.0 directly improves fruit quality; temperature range between 20 and 28°C, as excessively high or low temperatures may affect fruit quality or cause premature or delayed ripening; humidity range between 50% and 60%, as appropriately reducing humidity during the ripening period helps fruit ripening and sugar accumulation, while preventing diseases caused by excessive humidity; light intensity range between 2000 and 4000 lux, as sufficient but not excessive light is beneficial for photosynthesis and quality formation during fruit ripening; and pesticide residues must strictly comply with food safety standards, as excessive pesticide residues at the ripening stage will directly affect the quality and safety of agricultural products.
[0027] Furthermore, data from multi-source monitoring equipment used to monitor the farmland environment will be acquired periodically, including:
[0028] Based on the content of the first monitoring standard, a monitoring device that meets the first monitoring standard is configured, and the monitoring device includes a pH sensor;
[0029] Based on the content of the second monitoring standard, a monitoring device that meets the second monitoring standard is configured, and the monitoring device includes a temperature sensor, a humidity sensor and a light sensor;
[0030] Based on the content of the third monitoring standard, monitoring equipment that meets the third monitoring standard is configured, including a pesticide residue detector.
[0031] Based on the multi-stage growth of the crops, monitoring data from pH sensors, temperature sensors, humidity sensors, light sensors, and pesticide residue detectors are acquired periodically.
[0032] Specifically, pH sensors are evenly distributed throughout the farmland. For large areas of farmland, a grid layout is used to ensure accurate reflection of pH changes across the entire soil surface. The installation depth of the sensors is determined based on the root system distribution of the crops: 10-20 cm for shallow-rooted crops and 20-30 cm for deep-rooted crops. Temperature sensors are installed near the crop canopy to accurately measure the temperature of the growing environment. Humidity sensors are installed in well-ventilated locations that represent the humidity of the growing environment, avoiding direct sunlight and proximity to water sources. Light sensors are installed above the crop canopy, avoiding shading by leaves, to accurately measure the light intensity reaching the crop surface. Pesticide residue detection requires regular collection of soil, crop leaves, or fruit samples from the farmland, and testing according to the operating procedures of the testing instrument.
[0033] S200: Based on the multi-stage growth of the crop, data from multi-source monitoring devices for monitoring the farmland environment are periodically acquired, wherein the multi-source monitoring devices are communicatively connected to the intelligent control system.
[0034] Specifically, the data collection frequency is determined based on the crop's growth stage. During germination and seedling stages, when crops are more sensitive to environmental changes, the collection frequency is set at once every two hours. During the growth and maturity stages, the collection frequency is set at 1-2 times per day. Since pesticide residue testing requires manual sample collection, and the testing frequency depends on pesticide application, crop growth stage, and the safety standards for pesticide dosage at each stage, the collection frequency during germination and seedling stages is set at 14 days after pesticide application. During the growth and maturity stages, the testing frequency is increased to once every 7 days. After collecting various monitoring data, the data is transmitted to the intelligent control system using wireless communication technology.
[0035] Furthermore, before obtaining dimensionality-reduced multi-source monitoring equipment data, the following steps are also required:
[0036] Regularly acquire data from multi-source monitoring equipment for monitoring farmland environment, and perform data cleaning on the acquired multi-source monitoring equipment data. The data cleaning includes missing value handling, outlier detection, and standardization processing.
[0037] Specifically, the process involves checking each variable for missing values. For a small number of missing values, interpolation methods are used to fill in the gaps. If the proportion of missing values is large, statistical methods are used to estimate the missing values and fill them in. Outliers are detected and handled using statistical methods such as Z-score and interquartile range to identify them and correct or remove them. Formatting is standardized by using a unified date and time format to ensure all time information is stored in the same format. Numerical precision is standardized to ensure all numerical data is represented in the same way.
[0038] S300: The PCA algorithm is used to extract features and reduce the dimensionality of the acquired multi-source monitoring equipment data to obtain dimensionality-reduced multi-source monitoring equipment data, and the crop growth element set is updated. The crop growth element set is constructed based on the content of the multi-level farmland monitoring standard. The crop growth element set includes the soil quality element set, weather condition element set, and pesticide residue element set for crop growth.
[0039] Specifically, data acquired from multi-source monitoring devices is extracted and organized, including data collected from multiple pH sensors, temperature sensors, humidity sensors, light sensors, and pesticide residue detectors. A data matrix X is defined, where each row represents a sample, such as monitoring data from different time points or different farmland locations, and each column represents a feature, such as pH, temperature, or humidity. The covariance matrix Cov(X) of the data matrix X is calculated to understand the linear relationships between various features. The covariance matrix reflects the correlation between different features. If the covariance between two features is positive, it indicates a positive correlation; if it is negative, it indicates a negative correlation; and if it is close to 0, it indicates no correlation. The formula for the covariance matrix is:
[0040] Where n is the number of samples, The transpose of matrix X;
[0041] By solving the equation Cov(X)v=λV, we obtain the eigenvalues λ and eigenvectors V. The eigenvalues represent the magnitude of the variance explained by each principal component, while the eigenvectors determine the orientation of the principal components. The eigenvalues are arranged in descending order, and the corresponding eigenvectors are also sorted accordingly. Based on the set variance retention ratio or the number of principal components, the first k principal components are selected. Generally, principal components with a cumulative variance contribution rate reaching a certain proportion can be selected. For example, if the cumulative variance contribution rate of the first three principal components reaches 85%, these three principal components are selected for dimensionality reduction. The selection of principal components can determine which sensor features best reflect the key environmental conditions of this stage.
[0042] Dimensionality reduction is performed on the principal component data. For each sample Xi in the original data matrix X, dimensionality reduction is achieved by projecting it onto selected principal components. The dimensionality-reduced data is Yi = Xi * Vk, where Vk is a matrix composed of k selected eigenvectors.
[0043] S400: Set a sliding window with a sliding step size according to the multi-level growth stage of the crop. Calculate the average value of the data from each monitoring device updated in each element set of the crop growth element set within the window. Compare the average value of the data from each monitoring device in each element feature set with the corresponding monitoring standard in the multi-level farmland monitoring standard, calculate the deviation value, and obtain a multi-source deviation dataset.
[0044] Specifically, the sliding window size is set based on the time span of the multi-stage growth process and the frequency of data changes. If the growth stage is long and the data changes relatively slowly, a larger sliding window is used. Conversely, in stages of rapid growth, a smaller step size can be set to obtain more detailed data changes. For example, if the germination period is 10 days, a sliding window of 3 days and a sliding step size of 1 day would be used. In longer stages like the growth period, a sliding window of 5-7 days and a sliding step size of 2 days would be used.
[0045] Data from each monitoring device within the elemental feature set is acquired, and the average value is calculated within a sliding window. Taking the soil quality elemental set as an example, if there are multiple pH sensor data points within a sliding window, the sum of these data points is divided by the number of data points to obtain the average pH value within that sliding window. The same method is used to calculate the average value of monitoring device data for temperature, humidity, light intensity, pesticide residues, etc., within their respective elemental sets.
[0046] Obtain the multi-level farmland monitoring standards from step S100 and calculate the deviation values. Taking soil pH as an example, if the average pH value calculated within a certain sliding window is 6.5, while the standard range for that growth stage is 6.0-7.0, then the deviation value is 0. If the calculated average value is 5.5, the deviation value is -0.5. Using the same method, calculate the deviation values of each element, such as temperature, humidity, light intensity, and pesticide residues, from their corresponding standards. Combine these deviation values to form a multi-source deviation dataset. The multi-source deviation dataset reflects the deviation of the crop growth environment from the ideal monitoring standards within different sliding windows, providing an important data foundation for subsequent analysis.
[0047] S500: Based on the sliding step size, the multi-source deviation data in each window are arranged in order of the sliding step length to form a time-series multi-source deviation dataset.
[0048] Specifically, a multi-source deviation dataset is acquired, and a starting window is determined. The starting time of this window is the start time of the entire monitoring process. Starting from the first sliding window, the data is arranged sequentially according to the set sliding step size. For example, if the sliding step size is 1 day, and the data in the first window is the multi-source deviation data for days 1-3, then the data in the second window is arranged immediately after the data in the first window. For the multi-source deviation data in each sliding window, the order of the deviation values corresponding to each element (soil quality, weather conditions, pesticide residues, etc.) remains unchanged; only the data from different windows are arranged in chronological order. If there are n sliding windows, after this arrangement, a time-series multi-source deviation dataset containing n sets of multi-source deviation data is formed. This dataset reflects the changes of multi-source deviation data over time throughout the entire monitoring period.
[0049] S600: Input the time-series multi-source deviation dataset into the crop growth prediction model to obtain the crop yield prediction value, wherein the prediction model is built with ARIMAX as the architecture.
[0050] Further, crop growth prediction models include:
[0051] The historical annual yield of the crop is obtained and converted into multi-level growth stage data to obtain sample yield data for multi-level growth stages.
[0052] The sample soil quality element set, sample weather condition element set, and sample pesticide residue element set were obtained for multiple growth stages. The sample soil quality element set included pH value sensor sample data, the sample weather condition element set included temperature sensor sample data, humidity sensor sample data, and light sensor sample data, and the sample pesticide residue element set included pesticide residue detector sample data.
[0053] Set the sample sliding window and sample sliding step size to obtain the sample multi-source deviation dataset, and construct the sample time series multi-source deviation dataset. The sample multi-source deviation dataset includes pH sensor sample deviation data, temperature sensor sample deviation data, humidity sensor sample deviation data, light sensor sample deviation data, and pesticide residue detector sample deviation data.
[0054] By associating sample time-series multi-source deviation data with sample yield data of multi-stage growth, sample time-series data are extracted from the sample time-series multi-source deviation dataset to form a sample time-series dataset.
[0055] The sample time series dataset is used as the target variable, and the sample time series multi-source bias dataset is used as the exogenous variable. These are input into the ARIMAX model to determine the annual output of the sample and use it as the training sample.
[0056] Based on the training samples, the ARIMAX model is trained under supervision to generate a crop growth prediction model;
[0057] Based on the training samples, the crop growth prediction model is validated by performing model convergence checks and retraining until the convergence condition is met, thus obtaining the completed crop growth prediction model.
[0058] Furthermore, crop yield prediction methods based on multi-data fusion also include:
[0059] Regularly acquire data from multi-source monitoring equipment for monitoring farmland environment, and simultaneously obtain predicted crop yield data through crop growth prediction models;
[0060] The coordinate axis curves are determined, with the horizontal axis representing the time variable, which depends on the multi-stage growth of the crop. For the vertical axis, the acquired multi-source monitoring equipment data are plotted in different sub-graphs, and the multi-stage farmland monitoring standards and predicted yields corresponding to each monitoring equipment data are plotted in the corresponding sub-graphs, using dual coordinate axis labels.
[0061] Data visualization technology is used to display the drawn sub-graphs on the intelligent control system. Based on the data analysis and prediction results, a decision support system is built to provide decision support for improving crop yields.
[0062] Specifically, historical data from previous years are collected as samples to generate a sample time series dataset. Using the ARIMAX model, the sample time series dataset is used as the target variable to determine the sample annual output and create a prediction model.
[0063] Specifically, a time-series multi-source bias dataset is acquired and transmitted into the prediction model. The acquired multi-source data are then plotted in separate subplots. For example, a separate subplot is created for soil moisture data to visually demonstrate how soil moisture changes over time. The subplot uses dual axes: one axis represents the multi-source bias data (e.g., the left axis represents the soil moisture bias value), and the other axis represents the yield data corresponding to the bias value calculated by the prediction model. This allows for a clear display of different data types and their relationships within a single subplot, clearly defining the crop yield data under that condition.
[0064] Example 2
[0065] Based on the same inventive concept as the crop yield prediction method based on multi-data fusion in the foregoing embodiments, such as Figure 2 As shown, this disclosure also provides a crop yield prediction system based on multi-data fusion, the system comprising:
[0066] A multi-level farmland monitoring standard configuration module 11 is used to configure multi-level farmland monitoring standards for defined multi-level growth stages of crops, and to configure the multi-level farmland monitoring standards in an intelligent control system. The multi-level growth stages include germination period, seedling period, growth period and maturity period.
[0067] A multi-source monitoring equipment data acquisition module 12 is used to periodically acquire multi-source monitoring equipment data for monitoring the farmland environment based on the multi-level growth stages of the crop, wherein the multi-source monitoring equipment is communicatively connected to the intelligent control system.
[0068] The multi-source monitoring equipment data feature extraction module 13 is used to extract features and reduce dimensions of the acquired multi-source monitoring equipment data using the PCA algorithm to obtain dimensionality-reduced multi-source monitoring equipment data and update the crop growth element set. The crop growth element set is constructed based on the content of the multi-level farmland monitoring standard and includes the soil quality element set, weather condition element set, and pesticide residue element set for crop growth.
[0069] The multi-source deviation dataset acquisition module 14 is used to set a sliding window, the sliding step size is set according to the multi-level growth stage of the crop, calculate the average value of each monitoring device data updated in each element set of the crop growth element set within the window, compare the average value of each monitoring device data in each element feature set with the corresponding monitoring standard in the multi-level farmland monitoring standard, calculate the deviation value, and obtain the multi-source deviation dataset.
[0070] The time-series multi-source deviation dataset construction module 15 is used to arrange the multi-source deviation data in each window according to the sliding step length based on the sliding step length to form a time-series multi-source deviation dataset.
[0071] The crop yield prediction module 16 is used to input the time-series multi-source deviation dataset into the crop growth prediction model to obtain the crop yield prediction value, wherein the prediction model is built with ARIMAX as the architecture.
[0072] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0073] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this patent application should be determined by the appended claims.
Claims
1. A method for predicting crop yield based on multi-data fusion, characterized in that, The method includes: For the defined multi-level growth stages of crops, multi-level farmland monitoring standards are configured, and the multi-level farmland monitoring standards are configured in the intelligent control system. The multi-level growth stages include germination period, seedling period, growth period and maturity period. Based on the multi-stage growth of the crops, data from multi-source monitoring devices for monitoring the farmland environment are periodically acquired, wherein the multi-source monitoring devices are communicatively connected to the intelligent control system. The PCA algorithm is used to extract features and reduce the dimensionality of the acquired multi-source monitoring equipment data to obtain dimensionality-reduced multi-source monitoring equipment data, and update the crop growth element set. The crop growth element set is constructed based on the content of the multi-level farmland monitoring standard and includes the soil quality element set, weather condition element set, and pesticide residue element set for crop growth. Set a sliding window with a sliding step size according to the multi-level growth stage of the crop. Calculate the average value of the data from each monitoring device updated in each element set of the crop growth element set within the window. Compare the average value of the data from each monitoring device in each element feature set with the corresponding monitoring standard in the multi-level farmland monitoring standard, calculate the deviation value, and obtain a multi-source deviation dataset. Based on the sliding step size, the multi-source bias data in each window are arranged in order of the sliding step length to form a time-series multi-source bias dataset. The time-series multi-source bias dataset is input into the crop growth prediction model to obtain the crop yield prediction value, wherein the prediction model is built with ARIMAX as the architecture.
2. The crop yield prediction method based on multi-data fusion as described in claim 1, characterized in that, Before obtaining dimensionality-reduced multi-source monitoring equipment data, the following steps are required: Regularly acquire data from multi-source monitoring equipment for monitoring farmland environment, and perform data cleaning on the acquired multi-source monitoring equipment data. The data cleaning includes missing value handling, outlier detection, and standardization processing.
3. The crop yield prediction method based on multi-data fusion as described in claim 1, characterized in that, Multi-level farmland monitoring standards include: configuring a first monitoring standard, which is based on soil quality at multiple growth stages of crops; Configure a second monitoring standard, which is based on weather conditions at multiple growth stages of crops; Configure a third monitoring standard, which is based on pesticide residues at multiple growth stages of crops; By performing an index interaction correlation analysis on the first monitoring standard, the second monitoring standard, and the third monitoring standard, multi-level farmland monitoring standards are determined.
4. The crop yield prediction method based on multi-data fusion as described in claim 3, characterized in that, Regularly acquire data from multi-source monitoring equipment used to monitor the farmland environment, including: Based on the content of the first monitoring standard, a monitoring device that meets the first monitoring standard is configured, and the monitoring device includes a pH sensor; Based on the content of the second monitoring standard, a monitoring device that meets the second monitoring standard is configured, and the monitoring device includes a temperature sensor, a humidity sensor and a light sensor; Based on the content of the third monitoring standard, monitoring equipment that meets the third monitoring standard is configured, including a pesticide residue detector. Based on the multi-stage growth of the crops, monitoring data from pH sensors, temperature sensors, humidity sensors, light sensors, and pesticide residue detectors are acquired periodically.
5. The crop yield prediction method based on multi-data fusion as described in claim 4, characterized in that, Crop growth prediction models include: The historical annual yield of the crop is obtained and converted into multi-level growth stage data to obtain sample yield data for multi-level growth stages. The sample soil quality element set, sample weather condition element set, and sample pesticide residue element set were obtained for multiple growth stages. The sample soil quality element set included pH value sensor sample data, the sample weather condition element set included temperature sensor sample data, humidity sensor sample data, and light sensor sample data, and the sample pesticide residue element set included pesticide residue detector sample data. Set the sample sliding window and sample sliding step size to obtain the sample multi-source deviation dataset, and construct the sample time series multi-source deviation dataset. The sample multi-source deviation dataset includes pH sensor sample deviation data, temperature sensor sample deviation data, humidity sensor sample deviation data, light sensor sample deviation data, and pesticide residue detector sample deviation data. By associating sample time-series multi-source deviation data with sample yield data of multi-stage growth, sample time-series data are extracted from the sample time-series multi-source deviation dataset to form a sample time-series dataset. The sample time series dataset is used as the target variable, and the sample time series multi-source bias dataset is used as the exogenous variable. These are input into the ARIMAX model to determine the annual output of the sample and use it as the training sample. Based on the training samples, the ARIMAX model is trained under supervision to generate a crop growth prediction model; Based on the training samples, the crop growth prediction model is validated by performing model convergence checks and retraining until the convergence condition is met, thus obtaining the completed crop growth prediction model.
6. The crop yield prediction method based on multi-data fusion as described in claim 1, characterized in that, Also includes: Regularly acquire data from multi-source monitoring equipment for monitoring farmland environment, and simultaneously obtain predicted crop yield data through crop growth prediction models; The coordinate axis curves are determined, with the horizontal axis representing the time variable, which depends on the multi-stage growth of the crop. For the vertical axis, the acquired multi-source monitoring equipment data are plotted in different sub-graphs, and the multi-stage farmland monitoring standards and predicted yields corresponding to each monitoring equipment data are plotted in the corresponding sub-graphs, using dual coordinate axis labels. Data visualization technology is used to display the drawn sub-graphs on the intelligent control system. Based on the data analysis and prediction results, a decision support system is built to provide decision support for improving crop yields.
7. A crop yield prediction system based on multi-data fusion, characterized in that, The system includes: A multi-level farmland monitoring standard configuration module is used to configure multi-level farmland monitoring standards for defined multi-level growth stages of crops, and to configure the multi-level farmland monitoring standards in an intelligent control system. The multi-level growth stages include germination period, seedling period, growth period and maturity period. A multi-source monitoring equipment data acquisition module is used to periodically acquire multi-source monitoring equipment data of the farmland environment based on the multi-level growth stages of the crop, wherein the multi-source monitoring equipment is communicatively connected to the intelligent control system. A multi-source monitoring equipment data feature extraction module is used to extract features and reduce dimensions of the acquired multi-source monitoring equipment data using the PCA algorithm to obtain dimensionality-reduced multi-source monitoring equipment data and update the crop growth element set. The crop growth element set is constructed based on the content of the multi-level farmland monitoring standard and includes a soil quality element set, a weather condition element set, and a pesticide residue element set for crop growth. The multi-source deviation dataset acquisition module is used to set a sliding window with a sliding step size set according to the multi-level growth stage of the crop, calculate the average value of the monitoring equipment data updated in each element set of the crop growth element set within the window, compare the average value of the monitoring equipment data in each element feature set with the corresponding monitoring standard in the multi-level farmland monitoring standard, calculate the deviation value, and obtain the multi-source deviation dataset. A time-series multi-source deviation dataset construction module is used to arrange the multi-source deviation data in each window according to the sliding step length based on the sliding step length to form a time-series multi-source deviation dataset. The crop yield prediction module is used to input the time-series multi-source deviation dataset into the crop growth prediction model to obtain the crop yield prediction value, wherein the prediction model is built with ARIMAX as the architecture.
Citation Information
Patent Citations
Crop yield prediction algorithm based on multi-source heterogeneous agricultural data
CN117455062A
Crop management method and device based on large model, and medium
CN119027059A