A Machine Learning-Based Smart Agriculture Data Analysis Method and System
Through machine learning methods, agricultural data items are divided and classified, and compression strategies are dynamically adjusted, which solves the redundant storage and computing efficiency problems in agricultural data processing, and improves data management efficiency and decision-making reliability.
Patent Information
- Application Number
- CN202510608907.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-05-13
AI Technical Summary
The existing agricultural data processing methods are insufficient in redundant storage, computing efficiency bottlenecks and dynamic adaptability, and cannot effectively process agricultural timing data, resulting in waste of storage resources and calculation delays, and the compression strategy is out of touch with the demands of agricultural scenarios.
Smart agricultural data analysis method based on machine learning, divides data items through crop growth cycle, classifies them into similar and heterogeneous data items according to functions, dynamically adjusts compression and simplify strategies, combines abnormal detection and timing correlation analysis to optimize data processing flow.
Significantly reduce redundant storage, improve data management efficiency and decision-making reliability, provide accurate basis for agricultural operations, and reduce invalid calculations and loss of critical information.
Smart Images

Figure CN120123703B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a method and system for analyzing smart agriculture data based on machine learning. Background Art
[0002] In recent years, with the in-depth application of Internet of Things and sensor technologies in the agricultural field, time series data (such as soil temperature and humidity, fertilization amount, crop growth indicators) has shown exponential growth. However, existing data processing methods face the following core challenges in agricultural scenarios:
[0003] Redundant data storage burden: Traditional compression methods rely on fixed thresholds or static rules and cannot dynamically identify the similarity of adjacent data items (such as the same temperature value collected multiple times in a row), resulting in a large accumulation of duplicate data and significant waste of storage resources;
[0004] Computational efficiency bottleneck: Existing technologies uniformly process all data, ignoring time series characteristics (such as data fluctuation frequency, interval density differences). For example, when the sensor is normal, the data generation interval is stable, and when it is abnormal, the interruption is chaotic, but existing methods do not differentiate the processing, resulting in ineffective calculations and response delays;
[0005] Insufficient dynamic adaptability: General compression algorithms (such as Huffman coding) do not combine the functional attributes of data items (such as single-functional parameters and composite-functional parameters), and it is difficult to achieve efficient streamlining while retaining key features (such as the correlation between soil humidity and fertilization amount during the fertilization period), resulting in reduced interpretability of the compressed data or loss of key information.
[0006] Although existing research attempts to optimize the data volume through clustering or feature selection, the rule design lacks the joint application of dynamic analysis of time intervals and functional classification (such as separately processing single parameters and multi-parameter combinations), resulting in the disconnection between the compression strategy and the actual needs of the agricultural scenario. For example, it does not distinguish between the normal state of the sensor (high-frequency stable data can be compressed) and the abnormal state (low-frequency noise needs to be simplified), nor does it mine the mutation points at key growth stages through time series correlation (such as comparing data during the sowing period and the harvesting period). Summary of the Invention
[0007] In order to overcome the shortcoming of low efficiency in processing agricultural time series data, the present invention provides a method and system for analyzing smart agriculture data based on machine learning.
[0008] The technical implementation solution of the present invention is: A method for analyzing smart agriculture data based on machine learning, comprising the following steps:
[0009] S1: Collect agricultural analysis data, and construct a data sample of the agricultural analysis data; based on the data sample, divide data items according to the chronological order of crop growth to obtain a division result; based on the division result, obtain a complete data sample;
[0010] S2: Based on the data items, divide the data items into similar data items and dissimilar data items according to the functions of the data items; based on the data sample, extract the time interval data of the data items; based on the time interval data, perform a simplification process on the data items to obtain a simplification result;
[0011] S3: Based on the simplification result, obtain compressed data items and simplified data items; based on the compressed data items and the simplified data items, obtain a first analysis result of the agricultural analysis data; based on the first analysis result, extract the first data item and the last data item in the data sample;
[0012] S4: Based on the first data item and the last data item in the data sample, extract the first data item and the last data item in all data samples; based on the change conditions of the first data item and the last data item, then obtain a second analysis result of the agricultural analysis data.
[0013] Preferably, the collecting of the agricultural analysis data and the constructing of the data sample of the agricultural analysis data include:
[0014] The agricultural analysis data at least includes crop growth monitoring data, crop fertilization data, crop fertilization feedback data, and crop yield data during the crop growth process.
[0015] Preferably, the dividing of the data items according to the chronological order of crop growth based on the data sample to obtain a division result; and obtaining a complete data sample based on the division result includes:
[0016] Divide the data sample according to the chronological order of the crop growth process;
[0017] Based on the division result, obtain data items;
[0018] Based on the data items, obtain a data sample composed of N data items.
[0019] Preferably, the dividing of the data items into similar data items and dissimilar data items according to the functions of the data items includes:
[0020] Divide the data items in a data sample that implement specific functions and combined functions, and use the data items that implement specific functions as similar data items and the data items that implement combined functions as dissimilar data items.
[0021] Preferably, based on the data sample, extract the time interval data of the data item; based on the time interval data, perform a simplification process on the data item to obtain a simplification result, including:
[0022] Perform data standardization processing on the like data items and the unlike data items in the data sample;
[0023] Extract the time interval data between the like data items and the unlike data items, the time interval data between the like data items and the like data items, and the time interval data between the unlike data items and the unlike data items;
[0024] Use the length of the time interval data as the processing standard for the data item;
[0025] If the time interval is less than or equal to a preset time interval threshold, compress the like data item or the unlike data item;
[0026] If the time interval is greater than the preset time interval threshold, simplify the like data item or the unlike data item;
[0027] Based on the compression and the simplification process, obtain a simplification result.
[0028] Preferably, based on the simplification result, obtain a compressed data item and a simplified data item, including:
[0029] If the data item is processed by compression, delete the duplicate data in the data item, only retain the unique value, and retain the non-duplicate data intact, and define the data item formed by combining the data after the deletion process and the retention process as a compressed data item;
[0030] If the data item is processed by simplification, retain the same data in the data item, mark the importance according to the quantity distribution of the same data, and at the same time delete the data in the data item with a quantity less than a preset quantity threshold, and define the data item formed by combining the data after the deletion process and the retention process as a simplified data item.
[0031] Preferably, based on the compressed data item and the simplified data item, obtain a first analysis result of agricultural analysis data; based on the first analysis result, extract the first data item and the last data item in the data sample, including:
[0032] Based on the compressed data item and the simplified data item, obtain the quantities of the compressed data item and the simplified data item;
[0033] Use the quantity of the compressed data item and the quantity of all data items as a first ratio;
[0034] Take the number of the simplified data items and the number of all data items as a second ratio;
[0035] Obtain a first analysis result based on the first ratio and the second ratio;
[0036] The first data item includes at least one specific parameter in the crop growth monitoring data;
[0037] The last data item includes at least one type of data in the crop yield data.
[0038] Preferably, the extracting the first data item and the last data item in all data samples based on the first data item and the last data item in the data sample includes:
[0039] Construct a new data sample with all the first data items and define it as the first data sample;
[0040] Construct a new data sample with all the last data items and define it as the second data sample.
[0041] Preferably, the obtaining a second analysis result of the agricultural analysis data based on the change situation of the first data item and the last data item includes:
[0042] Obtain the data values of all data items in the second data sample, and starting from the first data item in the second data sample, using the outliers of the data values as splitting points, traverse the entire second data sample, taking the data items from the first data item to the first splitting point as one type of data items, taking the data items from the first splitting point to the second splitting point as the second type of data items, and so on, to obtain several types of data items;
[0043] Based on the positions of the splitting point data items in the second data sample, perform the same division on the first data sample; extract the compressed data items and the simplified data items in the first data sample, and use the ratio of the number of compressed data items to the number of data items in the corresponding type of data items and the ratio of the number of simplified data items to the number of data items in the corresponding type of data items to obtain a second analysis result of the agricultural analysis data.
[0044] Preferably, a smart agricultural data analysis system based on machine learning includes:
[0045] Data acquisition and sample construction module: Multisource collect crop growth monitoring data, fertilization data, fertilization feedback data and yield data through Internet of Things sensors, historical planting records, and manual inspection logs; Cut the original data into data items of continuous time periods according to the time axis of the crop growth cycle, and generate a structured data sample composed of N data items;
[0046] Data Classification and Simplification Processing Module: Classify according to the functional attributes of data items; perform standardized cleaning on similar and dissimilar data items, calculate the time intervals between adjacent data items, and perform compression or simplification operations based on preset thresholds;
[0047] Feature Compression and Preliminary Analysis Module: Remove duplicate data from the compressed data items, mark the importance of the simplified data items based on the data distribution density, and eliminate low-frequency data; count the proportions of the compressed and simplified data items, calculate the first ratio and the second ratio as key analysis indicators; extract the first and last items of the data sample to construct the first and last data indexes;
[0048] Time Series Comparison and In-depth Analysis Module: Aggregate the first items of all samples to construct the first data sample set, and the last items to construct the second data sample set; detect yield outliers in the second data sample set as segmentation points, and inversely map them to the first data sample set for synchronous segmentation; calculate the proportion differences of the compressed and simplified data items within each segment to generate the second analysis result.
[0049] Beneficial Effects: The present invention divides data items based on the chronological order of the crop growth cycle, and classifies data items into similar data items and dissimilar data items through functional attributes; after extracting the time interval data, dynamically adjusts the compression and simplification strategies, and optimizes the data processing flow according to the differences in sensor states. Through the first and last data comparison mechanism, combined with anomaly detection and segmentation analysis, effectively mines the time series correlation, reduces redundant storage while retaining key features, significantly improves the data management efficiency and decision-making reliability, and provides an accurate basis for optimizing agricultural operations. Description of the Drawings
[0050] Figure 1 It is a flowchart of the method for analyzing data of intelligent agriculture based on machine learning according to the present invention;
[0051] Figure 2 It is a structural diagram of the system for analyzing data of intelligent agriculture based on machine learning according to the present invention;
[0052] Figure 3 It is a schematic diagram of the process of constructing data samples according to the present invention;
[0053] Figure 4 It is a schematic diagram of the compression and simplification process according to the present invention;
[0054] Figure 5 It is a schematic diagram of the process of constructing the first data sample and the second data sample according to the present invention. Detailed Embodiments
[0055] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0056] Embodiment 1: A method for analyzing data of smart agriculture based on machine learning, as Figure 1 and Figures 3 - 5 shown, includes the following steps:
[0057] S1: Collect agricultural analysis data and construct a data sample of the agricultural analysis data; based on the data sample, divide data items according to the chronological order of crop growth to obtain a division result; based on the division result, obtain a complete data sample;
[0058] S2: Based on the data items, divide the data items into similar data items and dissimilar data items according to the functions of the data items; based on the data sample, extract the time interval data of the data items; based on the time interval data, perform simplification processing on the data items to obtain a simplification processing result;
[0059] S3: Based on the simplification processing result, obtain compressed data items and simplified data items; based on the compressed data items and the simplified data items, obtain a first analysis result of the agricultural analysis data; based on the first analysis result, extract the first data item and the last data item in the data sample;
[0060] S4: Based on the first data item and the last data item in the data sample, extract the first data item and the last data item in all data samples; based on the change conditions of the first data item and the last data item, and then obtain a second analysis result of the agricultural analysis data.
[0061] Collecting agricultural analysis data and constructing a data sample of the agricultural analysis data includes:
[0062] The agricultural analysis data at least includes crop growth monitoring data, crop fertilization data, crop fertilization feedback data, and crop yield data during the crop growth process.
[0063] It should be noted that modern agricultural planting relies on multi-source sensors to monitor crop growth and environmental parameters, and drives full-process automated operations, generating high-dimensional time-series data. In existing analysis methods, the long short-term memory (LSTM) model has the characteristics of a black box, manual experience is easily subject to subjective limitations, and the yield is affected by the coupling of multiple factors. It is necessary to systematically integrate growth parameters, environmental variables, and farm operation data to improve the reliability of decision-making.
[0064] Specifically, agricultural analysis data needs to be collected first. The specific process is as follows:
[0065] Before sowing, obtain the basic parameters of the plot through soil temperature and humidity sensors and light sensors; during the growth period: monitor growth indicators such as transpiration and leaf area index (LAI) in real time through crop physiological sensors (such as stem sap flow meters and leaf area meters); during the farming operation period: obtain the fertilization amount, fertilization time and feedback data based on the operation logs of Internet of Things devices (intelligent fertilizer applicators, drones) and manual inspection records; during the harvest period: collect the final yield data through yield sensors or manual entry into the system.
[0066] Thus, agricultural analysis data is obtained through the above method.
[0067] Based on the data sample, divide the data items according to the chronological order of crop growth to obtain the division result; based on the division result, obtain the complete data sample, including:
[0068] Divide the data sample according to the chronological order of the crop growth process.
[0069] Based on the division result, obtain the data items.
[0070] Based on the data items, obtain a data sample consisting of N data items.
[0071] It should be noted that as Figure 3 shown, the data sample constructed from agricultural analysis data includes various data throughout the entire growth cycle of the crop, including soil temperature and humidity, light intensity, fertilization amount, fertilization time, crop growth indicators (such as transpiration and leaf area index), and the final yield. Segment the data by growth stage (sowing period, growth period, fertilization period, maturity period), and combine the relevant data within each stage into "data items". For example: during the sowing period: combine "soil humidity" and "light intensity" into one data item; during the fertilization period: combine "soil humidity" and "fertilization amount and fertilization time" into one data item; during the maturity period: combine "leaf area index" and "transpiration" into one data item. Concatenate the data items of all stages in chronological order to form a new data sample. The original data is chaotic (for example, soil humidity, fertilization amount, and transpiration are scattered in different tables), and after segmentation, it is sorted by stage, clarifying the key parameter combinations for each stage. When analyzing the fertilization effect, directly look at the soil humidity and fertilization amount in the "fertilization period data item" without rummaging through other data; just process the corresponding data items by stage to reduce ineffective calculations.
[0072] Based on the data items, divide the data items into similar data items and dissimilar data items according to the function of the data items, including:
[0073] Divide the data items in a data sample that implement specific functions and combined functions. Use the data items that implement specific functions as homogeneous data items, and use the data items that implement combined functions as heterogeneous data items.
[0074] It should be noted that the data item classification method is as follows. Homogeneous data items only record data of a single function or parameter. For example: only record "soil humidity", or only record "light intensity". Heterogeneous data items: data that records multiple related functions or parameters at the same time. For example: record "soil humidity and fertilization amount" at the same time, or "light intensity and transpiration amount". Purpose of division: Homogeneous data items: used to quickly view the changes of a single parameter (such as only checking whether the soil humidity meets the standard); Heterogeneous data items: used to analyze how multiple parameters jointly affect the result (such as when the soil humidity is low, whether the fertilization amount needs to be adjusted). Example, use of homogeneous data items: quickly check whether a single parameter is abnormal (for example, trigger an irrigation alarm when the soil humidity is too low). Use of heterogeneous data items: analyze the relationship between parameters (for example, whether the transpiration amount increases when the light is strong, and whether the irrigation strategy needs to be adjusted).
[0075] It should be noted that a data sample consists of several data items, and the data items include homogeneous data items and heterogeneous data items. Among them, the quantities of homogeneous data items and heterogeneous data items are determined according to the actual situation.
[0076] Based on the data sample, extract the time interval data of the data items; based on the time interval data, perform a simplification process on the data items to obtain a simplification result, including:
[0077] Perform data standardization processing on the homogeneous data items and the heterogeneous data items in the data sample;
[0078] Extract the time interval data between the homogeneous data items and the heterogeneous data items, the time interval data between the homogeneous data items and the homogeneous data items, and the time interval data between the heterogeneous data items and the heterogeneous data items;
[0079] Use the length of the time interval data as the processing standard for the data items;
[0080] If the time interval is less than or equal to the preset time interval threshold, compress the homogeneous data items or the heterogeneous data items;
[0081] If the time interval is greater than the preset time interval threshold, simplify the homogeneous data items or the heterogeneous data items;
[0082] Based on the compression and the simplification process, obtain a simplification result.
[0083] It should be noted that for data standardization processing: Convert the data collected by different sensors into a unified format, including: Unit unification: Convert the temperature unit to °C and the humidity to a percentage; Data cleaning: Eliminate missing values and data outside the physical range (such as humidity > 100%); Time alignment: Unify the timestamp to millisecond-level accuracy.
[0084] Time interval extraction and threshold setting: Sensor response time: Record the time taken from detecting an anomaly (such as too high temperature) to triggering an operation (such as cooling down), and calculate the threshold. ; Data generation interval: Record the interval between two data transmissions of the sensor (such as once every 10 minutes), and calculate the threshold. ; Adjustment coefficient : Select according to business requirements (for example, in a loose scenario = 1, and in a strict scenario = 2).
[0085] State judgment and processing: If the response time ≤ and the generation interval ≤ , it is determined to be in a normal state, indicating that the sensor is working properly, the data is reliable, and duplicate data can be compressed (such as only retaining one of consecutive identical temperature values); If either interval exceeds the threshold, it is determined to be in an abnormal state, indicating that the sensor is faulty or has low efficiency, and the data needs to be simplified (such as deleting abnormal values that appear less than 2 times), and the alarm mechanism is triggered.
[0086] Among them, (mean response time): The average time taken by the sensor from detecting an anomaly (such as too high temperature) to triggering an operation (such as starting cooling down), (standard deviation of response time): Measure the volatility of the response time. The larger the standard deviation, the more unstable the response time. (mean generation interval): The average interval time between two data transmissions of the sensor, (standard deviation of generation interval): Measure the stability of the data generation interval. A small standard deviation indicates regular intervals, while a large one indicates interruptions or delays.
[0087] Based on the above simplification processing results, obtain compressed data items and simplified data items, including:
[0088] If the data item is for compression processing, delete the duplicate data in the data item, only retain the unique values, and keep the non-duplicate data intact. Define the data item formed by combining the data after deletion processing and retention processing as a compressed data item;
[0089] If the data item is to be simplified, the same data in the data item is retained, and importance marking is performed according to the quantity distribution of the same data. At the same time, the data with a quantity less than the preset quantity threshold in the data item is deleted, and the data item formed by combining the data after the deletion process and the retention process is defined as the simplified data item.
[0090] It should be noted that as Figure 4 shown, the compression processing steps are as follows: Delete duplicate data: If there are multiple identical data in the data item, only one is retained. Retain different data: All different data is retained. Form the compressed data item: Combine the retained data into a new data item. Purpose: Reduce redundant data and save storage space. Example: Original data item: 12°C, 12°C, 13°C, 13°C, 15°C, after compression: 12°C, 13°C, 15°C (delete the duplicate 12°C and 13°C). Example: When the sensor is working normally, the same temperature is collected continuously multiple times (such as in a constant temperature greenhouse), and only one valid data needs to be stored after compression.
[0091] The simplification processing steps are as follows: Mark importance: Count the number of times the same data appears, and the data with a higher number of appearances is marked as important data. And evaluate the influence weight of different parameters in the simplified data item on the yield through random forest to achieve intelligent marking. Machine learning can adapt to complex agricultural scenarios and improve the accuracy and generalization ability of data analysis. Delete low-frequency data: Delete the data with a number of appearances less than the preset quantity threshold (such as the number of appearances ≤ 2 times). Form the simplified data item: Combine the retained important data and the different data that have not been deleted into a new data item. Purpose: Eliminate noise data (such as occasional error values caused by sensor abnormalities) and retain key information. Example: Original data item: 12°C, 12°C, 13°C, 13°C, 15°C (assuming the preset quantity threshold = 2 times), after simplification: 12°C, 12°C, 13°C, 13°C (delete 15°C). Application scenario: When the sensor is abnormal, an abnormal temperature value (such as 15°C) is occasionally collected, and the noise is eliminated after simplification, and the main data is retained.
[0092] Based on the compressed data item and the simplified data item, obtain the first analysis result of the agricultural analysis data; based on the first analysis result, extract the first data item and the last data item in the data sample, including:
[0093] Based on the compressed data item and the simplified data item, obtain the quantities of the compressed data item and the simplified data item;
[0094] Take the quantity of the compressed data item and the quantity of all data items as the first ratio;
[0095] Take the quantity of the simplified data item and the quantity of all data items as the second ratio;
[0096] Obtain a first analysis result based on the first ratio and the second ratio;
[0097] The first data item includes at least one specific parameter in the crop growth monitoring data;
[0098] The last data item includes at least one type of data in the crop yield data.
[0099] It should be noted that the definition of the first ratio (the proportion of compressed data items): the number of compressed data items accounts for the proportion of all data items. The larger the first ratio: it indicates that there is more duplicate content in the data, the sensor works stably but there is a large amount of redundant data. For example: the temperature sensor sends the same data every hour (such as 25°C continuously 10 times), and only 1 time is retained after compression, and the proportion of compressed data items is high. Practical significance: reflects the degree of data redundancy, and when the ratio is high, it is necessary to optimize the acquisition frequency or storage strategy.
[0100] The definition of the second ratio (the proportion of simplified data items): the number of simplified data items accounts for the proportion of all data items. The larger the second ratio: it indicates that there is more abnormal or noisy data in the data, the sensor has a fault or there is a large environmental interference. For example: the humidity sensor frequently collects abnormal values (such as 40%, 80%, 45%), and the low-frequency abnormal values (such as 40%) are deleted after simplification, and the proportion of simplified data items is high. Practical significance: reflects the data noise level, and when the ratio is high, it is necessary to repair the equipment or optimize the anti-interference algorithm. Practical application analysis scenario 1: The first ratio is 80% and the second ratio is 5%. Conclusion: The data redundancy is serious but the noise is small. It is recommended to reduce the sensor sampling frequency (such as from once a minute to once every 10 minutes). Scenario 2: The first ratio is 20% and the second ratio is 50%. Conclusion: The data redundancy is small but the noise is large. It is necessary to check whether the sensor is damaged or the environmental interference source (such as electromagnetic interference). The significance of correlating the first and last data items. The first data item (such as the soil humidity at the sowing stage): represents the initial state. If its compression / simplification ratio is abnormal, it will affect subsequent operations; the last data item (such as the yield at the harvest stage): represents the final result. If the associated simplification ratio is high, it indicates that there is noise in the data during the critical growth period, resulting in a deviation in yield prediction.
[0101] Based on the first data item and the last data item in the data sample, extract the first data item and the last data item in all data samples, including:
[0102] Construct a new data sample with all the first data items and define it as the first data sample;
[0103] Construct a new data sample with all the last data items and define it as the second data sample.
[0104] It should be noted that as Figure 5As shown, the first data sample: Take out the first data item in all data samples (such as soil humidity and light intensity at sowing time) separately to form a new data set. The second data sample: Take out the last data item in all data samples (such as yield data at harvest) separately to form another new data set. Purpose: The first data sample: Analyze the impact of the initial conditions of crops (such as soil conditions) on overall growth; The second data sample: Analyze the correlation between the final yield and operations at each stage (such as fertilization and irrigation). Similarly, the correlation between all data samples constructed from all the second data items and all data samples constructed from all the last data items can be analyzed, and so on, which will not be elaborated here.
[0105] Based on the changes of the first data item and the last data item, and then obtain the second analysis result of agricultural analysis data, including:
[0106] Obtain the data values of all data items in the second data sample, and starting from the first data item in the second data sample, use the outliers of the data values as the splitting points to traverse the entire second data sample. Take the data items from the first data item to the first splitting point as one type of data item, and the data items from the first splitting point to the second splitting point as the second type of data item, and so on, to obtain several types of data items;
[0107] Based on the positions of the splitting point data items in the second data sample, make the same division for the first data sample; Extract the compressed data items and simplified data items in the first data sample, and use the ratio of the number of compressed data items to the number of data items in the corresponding type of data items and the ratio of the number of simplified data items to the number of data items in the corresponding type of data items, and then obtain the second analysis result of agricultural analysis data.
[0108] It should be noted that to find the production outliers: in the second data sample (production data), find the production values that are higher or lower than the preset production threshold as the splitting points. For example, the normal production is 400 - 600 kg / mu. If the production of a certain plot is 800 kg or 200 kg, it is marked as an abnormal splitting point. Split the production data: divide the production data into several segments according to the splitting points. Example: Splitting point 1: 800 kg; Splitting point 2: 200 kg; Segmentation result: high-yield segment (800 kg / mu), normal segment (400 - 600 kg / mu), low-yield segment (200 kg / mu). Synchronously split the sowing data: according to the position of the production splitting points, divide the first data sample (sowing date data, such as soil humidity, light) into corresponding segments. Example: Sowing data corresponding to the high-yield segment: soil humidity = 70%, light = 2200 Lx; Sowing data corresponding to the low-yield segment: soil humidity = 50%, light = 1500 Lx. Analyze the compression and simplification ratios: count the proportions of the compressed data items (after duplicate data deletion) and the simplified data items (after noise data deletion) in each segment. For example: High-yield segment: compression ratio 80% (data is stable), simplification ratio 5% (less noise); Low-yield segment: compression ratio 30% (data is unstable), simplification ratio 60% (more noise).
[0109] The first analysis result directly guides the data storage strategy by quantifying the proportions of compressed and simplified data; the second analysis result mines the correlation between abnormal growth stages and production through head-to-tail data comparison. The two form a technical closed-loop of "data optimization - precise decision-making - feedback correction", which not only improves the storage efficiency but also enhances the pertinence of agricultural operations. For example, if the soil humidity during the sowing period of a certain plot is abnormal (the first analysis result shows a high simplification ratio), resulting in a low production during the harvest period (the second analysis result locates the low-yield segment), an irrigation adjustment suggestion is automatically triggered, and at the same time, a sensor maintenance task is marked.
[0110] Example 2: Based on Example 1, a smart agriculture data analysis system based on machine learning, as Figure 2 shown, includes:
[0111] Data acquisition and sample construction module: Collect multi-source crop growth monitoring data, fertilization data, fertilization feedback data, and production data through Internet of Things sensors, historical planting records, and manual inspection logs; Cut the original data into continuous time period data items according to the time axis of the crop growth cycle, and generate a structured data sample consisting of N data items;
[0112] Data classification and simplification processing module: Complete classification according to the functional attributes of the data items; Perform standardized cleaning on similar and dissimilar data items, calculate the time interval between adjacent data items, and perform compression or simplification operations based on preset thresholds;
[0113] Feature Compression and Preliminary Analysis Module: Perform deduplication on the compressed data items, mark the importance of the simplified data items based on the data distribution density, and eliminate low-frequency data; count the proportion of the compressed and simplified data items, calculate the first ratio and the second ratio as the key analysis indicators; extract the first and last items of the data sample to construct the first and last data indexes.
[0114] Time Series Comparison and In-depth Analysis Module: Aggregate the first items of all samples to construct the first data sample set, and the last items to construct the second data sample set; detect the production outliers in the second data sample set as the segmentation points, and reverse map them to the first data sample set for synchronous segmentation; calculate the proportion difference of the compressed and simplified data items within each segment to generate the second analysis result.
[0115] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A machine learning-based method for analyzing data in smart agriculture, characterized in that Including the following steps: S1: Collect agricultural analysis data and construct a data sample of the agricultural analysis data; Based on the data sample, divide data items according to the chronological order of crop growth to obtain a division result; based on the division result, obtain a complete data sample; S2: Based on the data items, divide the data items into similar data items and dissimilar data items according to the functions of the data items; Based on the data sample, extract the time interval data of the data items; Based on the time interval data, perform a simplification process on the data items to obtain a simplification result, including: if the time interval is less than or equal to a preset time interval threshold, compress the similar data items or the dissimilar data items; delete duplicate data in the data items, only retain unique values, and define the data items formed by combining the data after the deletion process and the retention process as compressed data items; if the time interval is greater than the preset time interval threshold, simplify the similar data items or the dissimilar data items; retain the same data in the data items, mark the importance according to the quantity distribution of the same data, and at the same time delete the data items with the quantity less than a preset quantity threshold, and define the data items formed by combining the data after the deletion process and the retention process as simplified data items; S3: Based on the simplification result, obtain compressed data items and simplified data items; based on the compressed data items and the simplified data items, obtain a first analysis result of the agricultural analysis data; based on the first analysis result, extract the first data item and the last data item in the data sample; the first data item includes at least one specific parameter in the crop growth monitoring data; the last data item includes at least one type of data in the crop yield data; S4: Based on the first data item and the last data item in the data sample, extract the first data item and the last data item in all data samples; based on the change situation of the first data item and the last data item, then obtain a second analysis result of the agricultural analysis data, including: obtain the data values of all data items in the second data sample, and use the first data item in the second data sample as the starting point, use the outliers of the data values as the segmentation points, traverse the entire second data sample, use the data items from the first data item to the first segmentation point as one type of data items, use the data items from the first segmentation point to the second segmentation point as the second type of data items, and so on, to obtain several types of data items; based on the positions of the segmentation point data items in the second data sample, perform the same division on the first data sample; extract the compressed data items and the simplified data items in the first data sample, and use the ratio of the quantity of the compressed data items to the quantity of the data items in the corresponding type of data items and the ratio of the quantity of the simplified data items to the quantity of the data items in the corresponding type of data items, and then obtain the second analysis result of the agricultural analysis data.
2. The method for analyzing intelligent agriculture data based on machine learning according to claim 1 is characterized in that, The collection of agricultural analysis data and the construction of the data sample of the agricultural analysis data include: The agricultural analysis data at least includes crop growth monitoring data, crop fertilization data, crop fertilization feedback data, and crop yield data during the crop growth process.
3. A method for analyzing wisdom agriculture data based on machine learning according to claim 1, characterized in that, Based on the data sample, divide the data items according to the chronological order of crop growth to obtain a division result; Based on the division result, obtain a complete data sample, including: Divide the data sample in the chronological order of the crop growth process; Based on the division result, obtain data items; Based on the data items, obtain a data sample composed of N data items.
4. A method for analyzing intelligent agriculture data based on machine learning according to claim 1, characterized in that, Based on the data items, divide the data items into similar data items and dissimilar data items according to the functions of the data items, including: Divide the data items that implement specific functions and combined functions in a data sample. The data items that implement specific functions are used as similar data items, and the data items that implement combined functions are used as dissimilar data items.
5. A method for analyzing wisdom agriculture data based on machine learning according to claim 1, characterized in that, Based on the data sample, extract the time interval data of the data items; Based on the time interval data, perform a simplification process on the data items to obtain a simplification result, including: Perform data standardization processing on the similar data items and the dissimilar data items in the data sample; Extract the time interval data between the similar data items and the dissimilar data items, the time interval data between the similar data items and the similar data items, and the time interval data between the dissimilar data items and the dissimilar data items; Use the length of the time interval data as the processing standard for the data items; Based on the compression and the simplification process, obtain a simplification result.
6. A method for analyzing data in intelligent agriculture based on machine learning according to claim 1, characterized in that, Based on the compressed data items and the simplified data items, obtain a first analysis result of the agricultural analysis data; Based on the first analysis result, extract the first data item and the last data item in the data sample, including: Based on the compressed data items and the simplified data items, obtain the quantities of the compressed data items and the simplified data items; Use the quantity of the compressed data items and the quantity of all data items as a first ratio; Use the quantity of the simplified data items and the quantity of all data items as a second ratio; Based on the first ratio and the second ratio, obtain a first analysis result.
7. A method for analyzing data in smart agriculture based on machine learning according to claim 1, characterized in that, Based on the first data item and the last data item in the data sample, extract the first data item and the last data item in all data samples, including: Construct a new data sample with all the first data items and define it as the first data sample; Construct a new data sample with all the last data items and define it as the second data sample.
8. A machine learning-based intelligent agriculture data analysis system for implementing the machine learning-based intelligent agriculture data analysis method according to any one of claims 1-7, characterized in that, Including: Data acquisition and sample construction module: Collect crop growth monitoring data, fertilization data, fertilization feedback data, and yield data from multiple sources such as Internet of Things sensors, historical planting records, and manual inspection logs; Cut the original data into data items of continuous time periods according to the time axis of the crop growth cycle, and generate a structured data sample composed of N data items; Data classification and simplification processing module: Complete classification according to the functional attributes of data items; Perform standardization cleaning on similar and dissimilar data items, calculate the time interval between adjacent data items, and perform compression or simplification operations based on a preset threshold; Feature Compression and Preliminary Analysis Module: Perform deduplication on the compressed data items, mark the importance of the simplified data items based on the data distribution density, and eliminate low-frequency data; Statistically calculate the proportion of the compressed and simplified data items, calculate the first ratio and the second ratio as key analysis indicators; Extract the first and last items of the data sample to construct the first and last data indexes. Time Series Comparison and In-depth Analysis Module: Aggregate the first items of all samples to construct the first data sample set, and the last items to construct the second data sample set; Detect the production outliers in the second data sample set as the segmentation points, and inversely map them to the first data sample set for synchronous segmentation; Calculate the proportion difference of the compressed and simplified data items within each segment to generate the second analysis result.
Citation Information
Patent Citations
Agricultural non-point source pollution multi-source heterogeneous big data association method and big data supervision platform adopting same
CN110196886A
Agricultural park intelligent inspection system and method based on digital twinning
CN119006202A