Intelligent Detection and Filling Method for Missing Data Values in Agricultural Internet of Things
By obtaining the missing pattern features after dimensionality reduction in the agricultural Internet of Things and using multi-output neural networks for collaborative filling, the systematic lack of multi-source heterogeneous data is solved, the accuracy and consistency of data is achieved, adapted to changes in the dynamic environment, and the efficiency of data fusion is improved.
Patent Information
- Application Number
- CN202510459382.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-14
AI Technical Summary
The existing missing value detection and filling methods are difficult to effectively identify and process the differences in spatiotemporal synchronization of multi-source heterogeneous data in agricultural Internet of Things and systematic lacks caused by dynamic environments, resulting in difficult to guarantee data integrity and accuracy.
By obtaining the missing pattern characteristics after dimensionality reduction, mining the correlation rules between key dynamic factors and missing patterns, using multi-output neural networks for collaborative filling, and using consistency checking and adaptive learning rate adjustment strategies to ensure the accuracy and consistency of the fill data.
It improves the accuracy and consistency of the missing value filling process, adapts to dynamic environmental changes, reduces manual intervention, and improves the efficiency and accuracy of data fusion.
Smart Images

Figure CN119988850B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of missing value filling, and specifically to an intelligent detection and filling method for missing data in agricultural Internet of Things. Background Art
[0002] With the rapid development of Internet of Things (IoT) technology, various sensors and data acquisition devices have been gradually introduced in the agricultural field to achieve precision agricultural management and intelligent decision-making support. An agricultural IoT system usually integrates multiple sensors, such as soil moisture sensors, air temperature sensors, light intensity sensors, etc., and combines multi-source data sources, including satellite remote sensing images, data collected by drones, and data generated by ground monitoring devices. These multi-source heterogeneous data have different sampling frequencies and resolutions in the time and space dimensions, and provide comprehensive and detailed information support for crop growth environment monitoring, pest and disease warning, irrigation management, etc. through data fusion technology. However, the complexity and dynamics of the agricultural environment inevitably lead to data missing problems in the data collection process, affecting the integrity of the data and the accuracy of subsequent analysis.
[0003] In the Chinese invention patent with the application publication number CN116049672A, a method and device for filling missing data are provided to achieve accurate filling of missing values. The filling method includes: pre-filling the missing data to obtain pre-filled missing data, determining interpolation data and interpolation training data corresponding to the pre-filled missing data according to the pre-filled missing data, generating an adversarial network model, dividing the interpolation training data according to a preset length to obtain interpolation training vectors, training the adversarial network model according to the interpolation training vectors, inputting the interpolation data into the trained adversarial network model to obtain preliminary filled interpolation data, determining weight values corresponding to the preliminary filled interpolation data, and determining the filling value corresponding to the missing data according to the preliminary filled interpolation data and the weight values corresponding to the preliminary filled interpolation data.
[0004] In practical applications, the agricultural environment is highly dynamic. For example, weather changes and different crop growth stages lead to frequent failures of sensor devices, resulting in systematic data loss. At the same time, there are significant differences in time synchronization and mismatches in spatial resolution between different data sources. For instance, the update frequency of satellite imagery is relatively low, while the data acquisition frequency of soil moisture sensors is relatively high; the spatial resolution of drone data is relatively high, while the spatial coverage of satellite data is relatively wide. These factors jointly lead to the complexity and diversity of data loss patterns, including both random loss and systematic loss that depends on specific time periods or spatial regions. Existing methods for detecting and filling missing data values are insufficient in dealing with the spatio-temporal synchronization differences of multi-source heterogeneous data and systematic loss in a dynamic environment, making it difficult to effectively ensure the integrity and accuracy of data in the agricultural Internet of Things system, restricting the performance improvement and application promotion of intelligent agricultural decision-making systems.
[0005] In the agricultural Internet of Things, the spatio-temporal synchronization differences of multi-source heterogeneous data and the highly dynamic agricultural environment jointly result in complex data loss patterns, specifically manifested as systematic loss caused by frequent sensor failures. This loss not only has specific time and space patterns but is also highly correlated with environmental changes, resulting in non-random and dependent data loss.
[0006] Existing methods for detecting and filling missing values are difficult to effectively identify and handle this complex systematic loss, especially in the process of multi-source data fusion, where the spatio-temporal differences of different data sources and the dynamic changes of loss patterns cannot be fully considered, making it difficult to guarantee the accuracy and consistency of filling results. Therefore, there is an urgent need to develop a method that can intelligently detect and fill systematic missing values in the agricultural Internet of Things caused by spatio-temporal synchronization differences of multi-source heterogeneous data and dynamic environment, so as to improve data integrity and support intelligent decision-making for precision agriculture.
[0007] For this reason, the present invention provides an intelligent detection and filling method for missing data values in the agricultural Internet of Things. Summary of the Invention
[0008] (I) Technical problems to be solved
[0009] In view of the deficiencies of the prior art, the present invention provides an intelligent detection and filling method for missing data values in agricultural Internet of Things. By obtaining the missing pattern features after dimensionality reduction, mining the association rules between key dynamic factors and missing patterns, the current missing patterns are classified into different categories; after detecting and obtaining the missing categories of multi-source agricultural data, evaluating the correlation between different missing values, if the correlation exceeds the expectation, for highly correlated missing value groups, a multi-output neural network is used to predict multiple related missing values simultaneously, and a collaborative filling strategy is adopted; performing a consistency check on the filled data, if the consistency exceeds the expectation, seamlessly integrating the filled data with the original data, and transmitting the filled data to each data receiving end in real time; through the correlation-driven filling strategy, filling multiple associated missing values simultaneously, realizing overall filling, and improving the accuracy and consistency of the filling process; thus solving the technical problems recorded in the background art.
[0010] (II) Technical Solution
[0011] To achieve the above objectives, the present invention is realized through the following technical solutions: An intelligent detection and filling method for missing data values in agricultural Internet of Things, including, after collecting multi-source agricultural data by a sensor network, detecting, evaluating and marking abnormal data, and generating an abnormality degree from the obtained abnormal state data , if the abnormality degree exceeds the pre-abnormal threshold, sending an analysis instruction to the outside; among them, generating an abnormality degree from the abnormal state data, the method is as follows:
[0012] In the formula: and are weight coefficients, obtained by referring to the analytic hierarchy process, is the evaluation period is the number of abnormal data points within the evaluation period is the time is the severity of the abnormal data at time is the time is the location information of the abnormal data at time is the modulus length of the location information vector; is used to screen the time nodes where the severity of the abnormal data exceeds the severity threshold , ;
[0013] Performing dimensionality reduction on the multi-source data and obtaining the missing pattern features after dimensionality reduction, mining the association rules between key dynamic factors and missing patterns, and classifying the current missing patterns into different categories;
[0014] After detecting and obtaining missing classes of multi-source agricultural data, the correlation between different missing values is evaluated. If the correlation exceeds expectations, for highly correlated missing value groups, a multi-output neural network is used to simultaneously predict multiple related missing values, adopting a collaborative filling strategy;
[0015] After filling missing values, feedback data is collected. If the filling coefficient generated by the feedback data Below the filling threshold, an adaptive learning rate adjustment strategy is adopted;
[0016] Perform a consistency check on the filled data. If the consistency exceeds expectations, the filled data will be seamlessly integrated with the original data and transmitted to each data receiving end in real time. If the missing value filling effect is lower than expected, an alarm command will be issued to the outside.
[0017] Furthermore, data is collected by the sensor network, including soil moisture, temperature and light intensity data, as well as image data collected by drones, and summarized to generate a multi-source agricultural data set; the agricultural Internet of Things data is time-aligned, and the spatial resolution is unified using a multi-scale spatial interpolation method, and the spatial positions of different data sources are accurately aligned through a feature point-based registration algorithm.
[0018] Furthermore, the aligned agricultural data is used as input and the trained adaptive quality assessment model is used to perform quality assessment to automatically identify and mark potential abnormal data points; the abnormal status data of abnormal points in the agricultural data are identified, including the degree of abnormality and the time point when the abnormality occurred.
[0019] Furthermore, after receiving the analysis instruction, the preprocessed multi-source data are decomposed into intrinsic mode functions and a residual trend term using empirical mode decomposition, and the growth cycle of crops is divided into different stages in combination with the trained crop growth model, and relevant environmental variables are extracted as key dynamic factors in each growth stage.
[0020] Furthermore, a spatiotemporal correlation matrix is constructed, where represents the missing data at time point i and spatial location j, where the definition For binary variables:
[0021]
[0022] Applying t-distributed random neighborhood embedding to spatiotemporal correlation matrix Perform dimensionality reduction to obtain missing pattern features after dimensionality reduction;
[0023] Use the Apriori algorithm to mine the association rules between key dynamic factors and data missing, and describe how the key dynamic factors causally affect the data missing pattern through a pre-constructed Bayesian causal network. Based on the missing pattern features after dimensionality reduction, apply the pre-trained density clustering algorithm to classify the missing patterns into different categories.
[0024] Furthermore, take multi-source agricultural data as input, perform anomaly detection by the pre-trained graph neural network algorithm to obtain anomaly detection data, and use the pre-trained clustering algorithm to classify the anomaly detection data to obtain the corresponding missing classes;
[0025] According to the correspondence between the missing classes and the filling strategies, match the corresponding missing filling strategies for the corresponding missing classes from the filling strategy library to achieve targeted filling.
[0026] Furthermore, identify the dependency paths between missing values through the trained graph neural network in the following way: Define the relevance score to quantify the missing value and the missing value between the dependence strength;
[0027] The relevance score is calculated based on the edge weights in the graph model, , where, is the weight of the edge , indicating the missing value on the missing value influence strength; Construct the relevance matrix , where .
[0028] Furthermore, construct a filling model based on VAE. The filling model realizes generative filling of missing values by learning the latent distribution of the data. After obtaining the corresponding performance indicators after continuous times of filling, generate the filling coefficient in the following way. If the obtained filling coefficient is lower than the filling threshold, send a learning instruction to the outside, and generate the filling coefficient in the following way:
[0029]
[0030] Among them, represents the filling error metric, is the filling result of the th fold, is the true value; , are the weight coefficients, and the sum of the two is 1.
[0031] Further, after receiving the learning instruction, an adaptive learning rate adjustment strategy is adopted, and the Adam optimization algorithm is used to automatically adjust the learning rate to accelerate convergence and avoid falling into local optima, ensuring the continuous optimization of the model in a dynamic environment. An incremental learning method is adopted to gradually update the model parameters to adapt to newly arrived data.
[0032] Further, a multi-level consistency check mechanism is introduced. The spatio-temporal coherence constraint is used to ensure the consistency of the filled data in the time and space dimensions. If the consistency exceeds the expectation, the filled data is seamlessly integrated with the original data, and the filled data is transmitted to each data receiver in real time.
[0033] Monitor the performance metrics of the filling model. Use several key metrics as inputs, and use the feedback evaluation model after training to evaluate the current filling effect, and obtain the corresponding evaluation value. If the evaluation value does not exceed the preset evaluation threshold, collect the filling feedback data to optimize the filling model.
[0034] (III) Beneficial Effects
[0035] The present invention provides an intelligent detection and filling method for missing data in agricultural Internet of Things, which has the following beneficial effects:
[0036] 1. Through the standardization processing of the format and unit of the original data of the data source, and the application of time interpolation and space interpolation technologies, the consistency and compatibility of the data in the time and space dimensions are ensured; the problems of format, time and space inconsistency brought by multi-source data are eliminated, and the reliability and consistency of the data are ensured through advanced anomaly detection methods, reducing the complexity and computational burden of data fusion, and ensuring the effectiveness and accuracy of the entire missing value processing process.
[0037] 2. Identify the key dynamic factors affecting the agricultural environment, and deeply analyze the spatio-temporal distribution patterns of missing data by combining methods such as cluster analysis and principal component analysis, construct an association model between the dynamic environment and the missing pattern, and classify the missing pattern into random missing, time-related missing and space-related missing.
[0038] 3. Through dependency path identification and relevance scoring, when the relevance scoring is lower than expected, targeted filling methods are adopted for different types of missing values to achieve targeted processing; while when the relevance is higher, through the relevance-driven filling strategy, multiple associated missing values are filled simultaneously to achieve holistic filling, improving the accuracy and consistency of the filling process.
[0039] 4. After the current filling strategy fails to fully achieve efficient filling, ensure that the model can adapt to real-time environmental changes through dynamic parameter adjustment and incremental learning. Utilize new data and feedback information to continuously optimize the model parameters and structure, reduce manual intervention, and improve the efficiency and effectiveness of model optimization.
[0040] 5. Verify the accuracy and coherence of the filling results through consistency checking and multi-dimensional error evaluation. Especially in the data integration and feedback mechanism section, by integrating real-time data stream processing, geospatial indexing, and user feedback, the adaptability and filling effect of the system are significantly improved, ensuring the effectiveness and practicality of the filled data in actual applications. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 It is a schematic structural diagram of the intelligent detection and filling method for missing data values of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0043] Please refer to Figure 1 , the present invention provides an intelligent detection and filling method for missing data values for the agricultural Internet of Things, including
[0044] Step 1: After collecting multi-source agricultural data by the sensor network, detect and evaluate the data and mark the abnormal data, and generate an abnormality degree from the obtained abnormal state data. If the abnormality degree exceeds the pre-abnormal threshold, send an analysis instruction to the outside;
[0045] The content of the above Step 1 includes the following:
[0046] Step 101: Arrange a sensor network in the determined agricultural data collection area, including sensors such as soil humidity, air temperature, and light intensity. The sensor network collects data, including soil humidity, air temperature, and light intensity data, as well as image data collected by drones, etc., and aggregates them to generate a multi-source agricultural data set;
[0047] Perform time alignment processing on the agricultural Internet of Things data, including: using an adaptive time interpolation algorithm to downsample high-frequency data or upsample low-frequency data, and introducing a dynamic time window mechanism to dynamically adjust the size of the time window according to the update frequency and real-time requirements of the data source, ensuring the alignment of the data in the time dimension;
[0048] Step 102: For the spatial resolution differences of different data sources, a multi-scale spatial interpolation method is adopted to unify the spatial resolution. For example, for high-resolution data, downsampling is performed through fractal interpolation; for low-resolution data, upsampling is performed through a multi-scale interpolation method to ensure the consistency of all data in the spatial dimension; and different data sources are accurately aligned in spatial position through a feature-point-based registration algorithm to ensure consistency in the geographical space, so as to perform data fusion under the same geographical coordinate system;
[0049] During use, by standardizing the formats and units of the original data from different sensors (such as soil moisture, air temperature, light intensity) and data sources (such as satellite images, UAV data), and applying time interpolation and spatial interpolation techniques, the consistency and compatibility of the data in the time and spatial dimensions are ensured;
[0050] Step 103: Build an adaptive quality assessment model based on machine learning. Using the aligned agricultural data as input, the trained adaptive quality assessment model is used for quality assessment. By learning the quality characteristics (such as noise level, missing rate) of different data sources, potential abnormal data points are automatically identified and marked;
[0051] Identify the abnormal state data of the abnormal points in the agricultural data, including the degree of abnormality and the time node when the abnormality occurs, and generate an abnormality degree from the abnormal state data in the following way:
[0052]
[0053] In the formula: and are weight coefficients, obtained by referring to the analytic hierarchy process, is the evaluation period is the number of abnormal data points within the evaluation period is the time is the severity of the abnormal data at time is the time is the location information of the abnormal data at time is the modulus of the location information vector; is used to screen the time nodes when the severity of the abnormal data exceeds the severity threshold ; ;
[0054] According to historical data and management expectations for abnormal data, an abnormal threshold is set in advance; if the abnormality degree Exceeding the pre-set anomaly threshold indicates a high degree of anomaly in the current agricultural data and a large number of abnormal data, which requires targeted processing. For example, replacing abnormal data and filling missing values. At this time, an analysis instruction is sent to the outside.
[0055] When in use, combine the content in steps 101 to 103.
[0056] Through four links of data collection and standardization, time synchronization processing, spatial alignment and resolution unification, and data quality assessment and anomaly detection, the original data from multi-source heterogeneous data in the agricultural Internet of Things system has been comprehensively pre-processed and synchronized; it can eliminate the format, time, and spatial inconsistency problems brought by multi-source data, and also ensure the reliability and consistency of the data through advanced anomaly detection methods, reduce the complexity and computational burden of data fusion, and ensure the effectiveness and accuracy of the entire missing value processing process.
[0057] Step 2: Perform dimensionality reduction on multi-source data and obtain the missing pattern features after dimensionality reduction. After mining the association rules between key dynamic factors and missing patterns, classify the current missing patterns into different categories.
[0058] The said step 2 includes the following content:
[0059] Step 201: After receiving the analysis instruction, use empirical mode decomposition to decompose the pre-processed multi-source data into intrinsic mode functions and a residual trend term; for example, for soil moisture data , its decomposition form is:
[0060]
[0061] Among them, represents the i-th intrinsic mode function, is the residual trend term;
[0062] Combine the trained crop growth model to divide the growth cycle of crops into different stages, such as the germination stage, vegetative growth stage, flowering stage, and fruiting stage; extract relevant environmental variables as key dynamic factors in each growth stage, such as soil moisture H, temperature T, and light intensity L;
[0063] Step 202: Construct a spatio-temporal correlation matrix , where represents the data missing situation at time point i and spatial position j. Among them, define as a binary variable:
[0064]
[0065] By analyzing the matrix The structure of can reveal the distribution characteristics of missing values in time and space.
[0066] Applying t-distributed random neighborhood embedding to spatiotemporal correlation matrix Dimensionality reduction is performed to map the high-dimensional spatiotemporal missing pattern to a low-dimensional space to capture the nonlinear relationship and potential structure in the missing pattern and obtain the missing pattern characteristics after dimensionality reduction; the optimization objective function of t-distributed random neighborhood embedding for:
[0067]
[0068] in, and Represent the midpoints of high-latitude and low-latitude space respectively similarity;
[0069] Step 203: Use the Apriori algorithm to mine association rules between key dynamic factors and data missing, wherein a frequent item set L and its support supp(L) are defined, and a high-confidence association rule L⇒M is extracted from it, where M represents the missing pattern; and a pre-constructed Bayesian causal network is used to describe how key dynamic factors causally affect the data missing pattern;
[0070] Based on the missing pattern features after dimensionality reduction, the pre-trained density clustering algorithm is used to classify the missing patterns into different categories, such as random missing, time-related missing, and space-related missing. PCA analysis is used to extract the main components of the missing pattern, reduce the dimension and retain the main missing pattern features.
[0071] When using, combine the contents in steps 201 to 203:
[0072] By applying time series analysis and crop growth models (such as the CROP model) to identify key dynamic factors affecting the agricultural environment (such as seasonal changes and weather patterns), and combining cluster analysis and principal component analysis (PCA) to deeply analyze the spatiotemporal distribution patterns of missing data, an association model between the dynamic environment and missing patterns was constructed, and the missing patterns were classified into random missing, time-related missing, and space-related missing.
[0073] Identifying and classifying complex missing patterns can improve the accuracy and reliability of missing classification. At the same time, it can also describe the deep impact mechanism of environmental changes on data missingness and provide a basis for subsequent missing value detection and classification.
[0074] Step 3: After detecting missing classes of multi-source agricultural data, evaluate the correlation between different missing values. If the correlation exceeds expectations, for highly correlated missing value groups, use a multi-output neural network to simultaneously predict multiple related missing values, and adopt a collaborative filling strategy;
[0075] Step 3 includes the following content:
[0076] Step 301: Using multi-source agricultural data as input, perform anomaly detection by a pre-trained graph neural network algorithm to obtain anomaly detection data, and use a pre-trained clustering algorithm to classify the anomaly detection data to obtain corresponding missing classes;
[0077] Previously, summarize several missing value filling strategies to generate a filling strategy library. According to the correspondence between the missing classes and the filling strategies, match the corresponding missing filling strategies for the corresponding missing classes by the filling strategy library to achieve targeted filling;
[0078] It should be noted that: For random missing, the missing values have no specific time or space pattern, usually caused by occasional failures of sensors or short-term communication interruptions. For time-related missing, the missing values are associated with specific time periods, such as missing in a certain season or under specific climate conditions, reflecting the impact of the dynamic environment on data collection. For space-related missing, the missing values are concentrated in specific spatial regions, such as frequent failures of sensors in a certain plot, which may be due to environmental factors or equipment layout problems. For mixed missing (MixedMissing), the missing values are affected by both time and space factors, showing complex spatio-temporal dependencies.
[0079] Step 302: Identify the dependency paths between the missing values through the trained graph neural network, determine the propagation paths and influence scopes of the missing values. For example, the failure of a certain key sensor may cause the data of multiple related sensors to be missing. Evaluate the correlation between different missing values in the following way:
[0080] Define a correlation score to quantify the missing value and the missing value between the dependence strength;
[0081] The correlation score can be based on conditional probability or calculated based on the edge weights in the graph model: , or: ; where is the weight of edge , indicating the missing value on the missing value influence strength;
[0082] Construct a correlation matrix , where , used to describe the dependence relationship between multiple missing values. Through the matrix , high-correlation missing value groups can be identified to guide the selection of subsequent filling strategies;
[0083] If the correlation exceeds the expectation, based on the high correlation between missing values, for highly correlated groups of missing values, a multi-output neural network is used to predict multiple related missing values simultaneously, and a collaborative filling strategy is adopted to fill multiple related missing values simultaneously, leveraging the dependency relationship between them to improve the filling accuracy.
[0084] During use, combine the content in Steps 301 and 302:
[0085] Based on the graph-based anomaly detection method, accurately identify and classify missing values and outliers in the agricultural Internet of Things system, clarify the specific types and spatio-temporal correlations of missing values. Through dependency path identification and correlation scoring, when the correlation score is lower than the expectation, adopt targeted filling methods for different types of missing values to achieve targeted processing; while when the correlation is high, through a correlation-driven filling strategy, fill multiple related missing values simultaneously to achieve holistic filling, improving the accuracy and consistency of the filling process.
[0086] Step Four: After filling the missing values, collect feedback data. If the filling coefficient generated from the feedback data is lower than the filling threshold, adopt an adaptive learning rate adjustment strategy.
[0087] The said Step Four includes the following content:
[0088] Step 401: Construct a filling model based on VAE. The filling model realizes generative filling of missing values by learning the latent distribution of the data. After obtaining the corresponding performance metrics after continuous fillings, generate the filling coefficient in the following manner
[0089]
[0090] where, represents the filling error metric (such as the mean squared error MSE), is the filling result of the th fold, is the true value; , are weight coefficients, and the sum of the two is 1;
[0091] If the obtained filling coefficient is lower than the filling threshold, it indicates that the current filling effect fails to meet the expectation, and the current filling process needs to be optimized, and a learning instruction is sent to the outside.
[0092] Step 402: After receiving the learning instruction, adopt an adaptive learning rate adjustment strategy, use the Adam optimization algorithm, automatically adjust the learning rate to accelerate convergence and avoid falling into local optima, ensuring the continuous optimization of the model in a dynamic environment, in the following manner
[0093]
[0094] Among them, is the model parameter, is the initial learning rate, and are the bias correction values of the first-order and second-order momentum respectively, is a small constant to prevent division by zero;
[0095] Adopt an incremental learning method to gradually update the model parameters to adapt to the newly arrived data, ensuring the continuous optimization of the model in a dynamic environment; through the online learning algorithm, the model can be adjusted in real time to reflect the latest environmental changes and data patterns;
[0096] When in use, combine the content in steps 401 and 402:
[0097] After the current filling strategy cannot fully achieve efficient filling, ensure that the model can adapt to real-time environmental changes through dynamic parameter adjustment and incremental learning, and use new data and feedback information to continuously optimize the model parameters and structure; reduce manual intervention and improve the efficiency and effect of model optimization.
[0098] Step Five: Perform a consistency check on the filled data. If the consistency exceeds the expectation, seamlessly integrate the filled data with the original data and transmit the filled data to each data receiving end in real time. If the missing value filling effect is lower than the expectation, send an alarm instruction to the outside;
[0099] The said Step Five includes the following content:
[0100] Step 501: Introduce a multi-level consistency check mechanism and use spatio-temporal coherence constraints to ensure the consistency of the filled data in the time and space dimensions , ensuring the consistency of the filled value with its neighboring points in time before and after and in space, in the following way:
[0101]
[0102] Among them: is the data matrix after filling, where is the tolerance threshold, used to define the acceptance range of consistency, and represent the time and space indices respectively, represents the coherence function, ensuring that the filled value is consistent with its previous time point and the previous space point , and the specific definition is:
[0103]
[0104] If the consistency exceeds expectations, seamlessly integrate the filled data with the original data and transmit the filled data to each data receiving end in real time;
[0105] Step 502: Monitor the performance metrics of the filling model, including filling error and processing delay, etc. Train a convolutional neural network with the labeled sample data to obtain the trained feedback evaluation model;
[0106] Take several key metrics as inputs, use the trained feedback evaluation model to evaluate the current filling effect, and obtain the corresponding evaluation value. If the evaluation value does not exceed the preset evaluation threshold, it means that the current filling performance fails to meet expectations and the current filling model needs to be iteratively optimized. At this time, send an alarm instruction to the outside;
[0107] After receiving the alarm instruction, collect user feedback and filling feedback data in actual applications, and continuously optimize the filling model through the feedback data;
[0108] When in use, combine the content in Steps 501 and 502:
[0109] Verify the accuracy and coherence of the filling results through consistency checking and multi-dimensional error evaluation. Especially in the data integration and feedback mechanism part, through real-time data stream processing, geospatial indexing, and user feedback integration, the adaptability and filling effect of the system are significantly improved, ensuring the effectiveness and practicality of the filled data in actual applications.
[0110] The Analytic Hierarchy Process (AHP) is a systematic multi-criteria decision-making method aimed at simplifying the analysis process by constructing a hierarchical structure model to decompose complex decision-making problems into multiple levels and elements. AHP first decomposes the decision-making problem into goals, criteria (or standards) and their subordinate sub-criteria or alternative solutions, and then conducts pairwise comparisons through experts or decision-makers to evaluate the importance of each element relative to the upper-level element. The eigenvector method is used to calculate the weights of each element, and the consistency of the judgment matrix is checked through the Consistency Ratio (CR) to ensure the rationality and reliability of the weight allocation. AHP is widely used in fields such as resource allocation, strategic planning, project evaluation, and risk management. Its intuitive structured analysis and quantitative weight determination method help decision-makers make scientific and reasonable decision-making choices under multiple dimensions and multiple criteria.
[0111] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0112] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0113] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only for some logical function divisions, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.
[0114] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0115] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. An intelligent detection and filling method for missing data values in the agricultural Internet of Things, characterized in that: including After collecting multi-source agricultural data by the sensor network, the data is detected, evaluated, and abnormal data is marked. The abnormality degree is generated from the obtained abnormal status data , if the abnormality degree exceeds the pre-abnormal threshold, an analysis instruction is sent to the outside; among them, the abnormality degree is generated from the abnormal status data , the method is as follows: ; where: and are weight coefficients, is the evaluation period is the number of abnormal data points within the evaluation period, is time is the severity of the abnormal data at time is time is the location information of the abnormal data at time is the modulus of the location information vector; among them, data collection is carried out by a sensor network, including soil moisture, air temperature and light intensity data, and image data collected by an unmanned aerial vehicle, and a multi-source agricultural data set is generated by summarization; Taking the aligned agricultural data as input, using the trained adaptive quality assessment model for quality assessment, automatically identifying and marking potential abnormal data points; identifying the abnormal status data of the agricultural data abnormal points, including the degree of abnormality and the time node when the abnormality occurs For screening the severity of abnormal data Exceeding the severe threshold Time node ; Reducing the dimension of multi-source data and obtaining the missing pattern features after dimension reduction, classifying the current missing patterns into different categories after mining the association rules between key dynamic factors and the missing patterns After receiving the analysis instruction, using empirical mode decomposition to decompose the preprocessed multi-source data into intrinsic mode functions and a residual trend term, combining the trained crop growth model to divide the growth cycle of crops into different stages, and extracting relevant environmental variables as key dynamic factors at each growth stage After detecting and obtaining the missing classes of multi-source agricultural data, evaluating the correlation between different missing values. If the correlation exceeds the expectation, for the highly correlated missing value group, using a multi-output neural network to predict multiple related missing values simultaneously and adopting a collaborative filling strategy Collect feedback data after filling in missing values. If the filling coefficient generated from the feedback data is lower than the filling threshold, adopt an adaptive learning rate adjustment strategy; Performing a consistency check on the filled data. If the consistency exceeds the expectation, seamlessly integrating the filled data with the original data and transmitting the filled data to each data receiving end in real time. If the missing value filling effect is lower than the expectation, sending an alarm instruction to the outside 2. The intelligent detection and filling method for data missing values according to claim 1, characterized in that Performing time alignment processing on agricultural Internet of Things data, using a multi-scale spatial interpolation method to unify the spatial resolution, and accurately aligning the spatial positions of different data sources through a feature point-based registration algorithm 3. The intelligent detection and filling method for data missing values according to claim 2, characterized in that Construct a spatio-temporal correlation matrix , where represents the data missing situation at time point i and spatial position j, where it is defined that is a binary variable: ; Apply t-distributed Stochastic Neighbor Embedding to the spatio-temporal correlation matrix for dimensionality reduction to obtain the missing pattern features after dimensionality reduction; Using the Apriori algorithm to mine the association rules between key dynamic factors and data missing, and describing how key dynamic factors causally affect the data missing pattern through a pre-constructed Bayesian causal network. Based on the missing pattern features after dimension reduction, applying a pre-trained density clustering algorithm to classify the missing patterns into different categories 4. The intelligent detection and filling method for data missing values according to claim 3, characterized in that Taking multi-source agricultural data as input, performing anomaly detection by a pre-trained graph neural network algorithm to obtain anomaly detection data, and classifying the anomaly detection data by a pre-trained clustering algorithm to obtain the corresponding missing classes According to the correspondence between the missing classes and the filling strategies, matching the corresponding missing filling strategies for the corresponding missing classes by the filling strategy library to achieve targeted filling 5. The intelligent detection and filling method for data missing values according to claim 4, characterized in that Identifying the dependence path between missing values through a trained graph neural network in the following way Define the relevance score to quantify the missing values and the missing values between the dependence strengths; the relevance score is calculated based on the edge weights in the graph model, , where is the weight of the edge indicating the influence strength of the missing value on the missing value , and construct a relevance matrix , where .
6. The intelligent detection and filling method for data missing values according to claim 5, characterized in that Construct a filling model based on VAE. The filling model realizes generative filling of missing values by learning the latent distribution of data. After obtaining the corresponding performance metrics after consecutive fillings; Generate the filling coefficient in the following manner , if the obtained filling coefficient is lower than the filling threshold, send a learning instruction externally to generate the filling coefficient in the following manner: ; wherein, represents the filling error metric, is the filling result of the th fold, is the true value; , are the weight coefficients, and the sum of the two is 1.
7. The intelligent detection and filling method for data missing values according to claim 6, characterized in that After receiving the learning instruction, an adaptive learning rate adjustment strategy is adopted, and the Adam optimization algorithm is used to automatically adjust the learning rate to accelerate convergence and avoid falling into local optima, ensuring the continuous optimization of the model in a dynamic environment. An incremental learning method is adopted to gradually update the model parameters to adapt to the newly arrived data.
8. The intelligent detection and filling method for data missing values according to claim 7, characterized in that: A multi-level consistency check mechanism is introduced. The spatio-temporal coherence constraint is used to ensure the consistency of the filled data in the time and space dimensions. If the consistency exceeds the expectation, the filled data is seamlessly integrated with the original data, and the filled data is transmitted to each data receiving end in real time; Monitor the performance indicators of the filling model. Take several key indicators as inputs, and use the obtained feedback evaluation model after training to evaluate the current filling effect and obtain the corresponding evaluation value. If the evaluation value does not exceed the preset evaluation threshold, collect the filling feedback data to optimize the filling model.
Citation Information
Patent Citations
Missing data filling method and device
CN116049672A
Distribution network voltage data missing filling method and device
CN111309718A
Aero-engine parameter-oriented missing value filling method based on cross-neighborhood clustering and time sequence dependence
CN118296757A