Multi-source heterogeneous agricultural data real-time cleaning method and system

By employing a real-time cleaning method for multi-source heterogeneous agricultural data, and utilizing data fusion and a pre-trained cleaning network to unify data formats and remove noise, this approach addresses the shortcomings of traditional cleaning methods in processing multi-source heterogeneous data. It achieves efficient and automated data cleaning, ensuring data quality and information integrity.

CN120670727BActive Publication Date: 2026-04-14NANTONG BIPU TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANTONG BIPU TECHNOLOGY CO LTD
Filing Date
2025-06-10
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Traditional data cleaning methods are difficult to effectively handle the heterogeneity, timeliness, and dynamism of multi-source heterogeneous agricultural data, resulting in poor cleaning effects. Furthermore, the lack of support from agricultural production temporal and spatial factors and domain knowledge can easily lead to data distortion or information loss.

Method used

A real-time cleaning method for multi-source heterogeneous agricultural data is adopted. Through data fusion, spatiotemporal alignment processing and pre-trained cleaning network, primary cleaning, format unification and secondary cleaning are performed. Dynamic spatiotemporal grid and pre-trained cleaning network are used for iterative denoising, simulating the data noise diffusion and reverse diffusion process, decoding and reconstructing text, and realizing data standardization and denoising.

Benefits of technology

It significantly improves data cleaning effectiveness, reduces manual intervention, lowers costs, retains key information, provides more accurate data support, and provides an efficient data foundation for agricultural production.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120670727B_ABST
    Figure CN120670727B_ABST
Patent Text Reader

Abstract

The application discloses a multi-source heterogeneous agricultural data real-time cleaning method and system, belongs to the technical field of data cleaning, and specifically comprises the following steps: acquiring multi-source heterogeneous data, performing primary cleaning on the acquired multi-source heterogeneous data, obtaining the multi-source heterogeneous data after primary cleaning, converting the multi-source heterogeneous data after primary cleaning into a standardized data format, performing spatio-temporal alignment processing on the data based on a dynamic spatio-temporal grid, obtaining the processed multi-source heterogeneous data, pre-training a cleaning network, inputting the processed multi-source heterogeneous data into the pre-trained cleaning network for iterative cleaning, randomly selecting the multi-source heterogeneous data text after cleaning in the iterative process and mapping the multi-source heterogeneous data text to a latent space, simulating a data noise forward diffusion and reverse diffusion process, decoding and reconstructing the text, and obtaining the multi-source heterogeneous data after final cleaning; and the application improves the efficiency and accuracy of data cleaning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data cleaning technology, specifically a method and system for real-time cleaning of multi-source heterogeneous agricultural data. Background Technology

[0002] With the rapid development of Internet of Things (IoT) technology, the agricultural sector is gradually introducing various sensors, cameras, satellite remote sensing equipment and other hardware to achieve real-time collection of information such as crop growth, soil moisture, environmental temperature and humidity, and meteorological data.

[0003] In agricultural IoT application scenarios, data sources are diverse, mainly including: satellite remote sensing data, ground flood station observation data, historical disaster data, soil moisture and crop growth data, etc. Due to differences in acquisition equipment, sampling frequency, data format and data transmission method, these data exhibit heterogeneity, timeliness and dynamism. At the same time, due to environmental noise, sensor failure, data packet loss and other reasons, there are often noise, missing and inconsistencies in the data.

[0004] Traditional data cleaning methods mostly focus on offline data processing, relying on manual rules or simple statistical methods to correct and complete data. These methods have the following shortcomings: limited ability to process heterogeneous data; for cleaning needs of multi-source heterogeneous data, a single rule or model cannot simultaneously adapt to different data types and sources, resulting in inconsistent data cleaning effects; and a lack of environmental context and domain knowledge support. Conventional methods struggle to incorporate spatiotemporal factors and domain knowledge in agricultural production during the cleaning process, easily leading to data distortion or loss of key information. Therefore, a real-time cleaning method for multi-source heterogeneous agricultural data is urgently needed. Summary of the Invention

[0005] To address the shortcomings of existing technologies, this invention proposes a real-time cleaning method and system for multi-source heterogeneous agricultural data. By utilizing data fusion, spatiotemporal alignment processing, and a complete process from vectorization processing, noise simulation, back diffusion denoising to decoding and reconstruction, it not only solves the deficiencies of traditional data cleaning in terms of real-time performance, multi-source adaptability, and automation, but also significantly improves the cleaning effect of agricultural data.

[0006] To achieve the above objectives, the present invention provides the following technical solution:

[0007] Real-time cleaning methods for multi-source heterogeneous agricultural data include:

[0008] Acquire multi-source heterogeneous data, perform cleaning on the acquired multi-source heterogeneous data, and obtain cleaned multi-source heterogeneous data.

[0009] After cleaning, the multi-source heterogeneous data is converted into a standardized data format, and spatiotemporal alignment is performed on the data based on a dynamic spatiotemporal grid to obtain the processed multi-source heterogeneous data.

[0010] A pre-trained cleaning network is used to input the processed multi-source heterogeneous data into the pre-trained cleaning network for iterative cleaning. The text of the multi-source heterogeneous data after cleaning is randomly selected during the iteration process and mapped to the latent space. The forward and backward diffusion processes of data noise are simulated, and the text is decoded and reconstructed to obtain the final cleaned multi-source heterogeneous data.

[0011] Specifically, the process of converting multi-source heterogeneous data after initial cleaning into a standardized data format and performing spatiotemporal alignment processing on the data based on a dynamic spatiotemporal grid includes:

[0012] Feature parsing and field extraction are performed on multi-source heterogeneous data after one cleaning process, the mapping relationship between data is analyzed, and a data format mapping table is established.

[0013] Data is spatiotemporally aligned based on a dynamic spatiotemporal grid.

[0014] Multi-source heterogeneous data is reconstructed and managed using dynamic metadata to obtain processed multi-source heterogeneous data.

[0015] Specifically, the spatiotemporal alignment processing of data based on a dynamic spatiotemporal grid includes:

[0016] The entire region is divided according to a preset initial standard grid, a time axis granularity is established, and sensitive terrain areas and crop planting zones are identified.

[0017] Establish a UTC standard timeline, calibrate and convert timestamps in multi-source heterogeneous data after one cleaning, and mark data of different frequencies with a unified time granularity.

[0018] Different projected coordinate systems are unified into the WGS84 geographic coordinate system using the seven-parameter transformation method or coordinate transformation library, and spatial range identifiers are added to each multi-source heterogeneous data.

[0019] Specifically, the reconstruction of multi-source heterogeneous data and the management using dynamic metadata to obtain processed multi-source heterogeneous data include:

[0020] The spatiotemporally aligned multi-source heterogeneous data is constructed into a spatiotemporal four-dimensional tensor, with dimensions including: longitude, latitude, time, and data source type.

[0021] Create data traceability archives to record the original data format, acquisition device ID, and processing version number; and build a quality labeling system, including integrity labels and consistency labels.

[0022] Generate standard data packets, which are processed multi-source heterogeneous data, including: spatiotemporal four-dimensional tensors, quality labels, and traceability archives.

[0023] Specifically, the pre-trained cleaning network inputs the processed multi-source heterogeneous data into the pre-trained cleaning network for iterative cleaning. During the iteration process, it randomly selects the cleaned multi-source heterogeneous data text and maps it to the latent space, simulating the forward and backward diffusion processes of data noise, and decoding and reconstructing the text, including:

[0024] A pre-trained cleaning network is used to input the processed multi-source heterogeneous data into the pre-trained cleaning network for iterative denoising. The text of the multi-source heterogeneous data after cleaning during the iteration process is randomly selected and mapped to a semantic embedding space. The embedded vector is then normalized.

[0025] Construct a forward diffusion process and gradually inject random noise according to a preset noise schedule;

[0026] Construct a backdiffusion cleaning network to gradually restore clean text vectors from a noisy embedding space;

[0027] The decoder is used to map clean text vectors back to text sequences, and the text sequences are used as input to a pre-trained denoising network for iterative processing.

[0028] The process involves repeated random noise injection, restoration, and mapping iterations to construct multi-dimensional evaluation metrics, including reconstruction accuracy, syntactic correctness, and semantic consistency. The output of the pre-trained denoising network is then evaluated, and the data with the best evaluation results is selected as the final cleaned multi-source heterogeneous data.

[0029] Specifically, the construction of the backdiffusion cleaning network to progressively restore clean text vectors from a noisy embedding space includes:

[0030] Based on a lightweight pre-trained cleaning network, a back-diffusion cleaning network is constructed.

[0031] By setting conditional branches and using the processed multi-source heterogeneous data as conditional inputs, the backdiffusion cleaning network is assisted in restoring the text structure and semantics under different noise levels.

[0032] A clean text representation is gradually restored from a noisy embedding space using a backdiffusion process.

[0033] Specifically, the multi-source heterogeneous data includes: satellite remote sensing data, ground flood station observation data, historical disaster data, soil moisture and crop growth data.

[0034] A real-time cleaning system for multi-source heterogeneous agricultural data is used to implement the real-time cleaning method for multi-source heterogeneous agricultural data, including: a primary cleaning module, a spatiotemporal processing module, and a secondary cleaning module;

[0035] The first cleaning module is used to acquire multi-source heterogeneous data, perform a first cleaning on the acquired multi-source heterogeneous data, and obtain multi-source heterogeneous data after the first cleaning.

[0036] The spatiotemporal processing module is used to convert multi-source heterogeneous data after one cleaning into a standardized data format, and perform spatiotemporal alignment processing on the data based on a dynamic spatiotemporal grid to obtain processed multi-source heterogeneous data.

[0037] The secondary cleaning module is used to pre-train the cleaning network. The processed multi-source heterogeneous data is input into the pre-trained cleaning network for iterative cleaning. The text of the multi-source heterogeneous data after cleaning during the iteration process is randomly selected and mapped to the latent space to simulate the forward and backward diffusion process of data noise. The text is decoded and reconstructed to obtain the final cleaned multi-source heterogeneous data.

[0038] Specifically, the spatiotemporal processing module includes: a parsing unit, a spatiotemporal alignment processing unit, and a reconstruction management unit;

[0039] The parsing unit is used to perform feature parsing and field extraction on multi-source heterogeneous data after one cleaning, analyze the mapping relationship between data, and establish a data format mapping table.

[0040] The spatiotemporal alignment processing unit is used to perform alignment processing on multi-source heterogeneous data after one cleaning based on time and space.

[0041] The reconstruction management unit is used to construct a spatiotemporal four-dimensional tensor and a data traceability archive, and to generate a standard data package.

[0042] Specifically, the secondary cleaning module includes: a vectorization processing unit, a forward diffusion unit, a reverse diffusion cleaning unit, a decoding unit, and a cleaning evaluation unit;

[0043] The vectorization processing unit is used to input the processed multi-source heterogeneous data into a pre-trained cleaning network for iterative denoising, randomly select the cleaned multi-source heterogeneous data text during the iteration process, map it to a semantic embedding space, and normalize the embedded vector.

[0044] The forward diffusion unit is used to construct the forward diffusion process and gradually inject random noise according to a preset noise schedule.

[0045] The backdiffusion cleaning unit is used to construct a backdiffusion cleaning network to gradually restore clean text vectors from a noisy embedding space.

[0046] The decoding unit is used to map clean text vectors back to text sequences using a decoder, and to iterate the text sequences as input to a pre-trained denoising network.

[0047] The cleaning and evaluation unit is used to evaluate the output of the pre-trained denoising network and select the data with the best evaluation result as the final cleaned multi-source heterogeneous data.

[0048] Compared with the prior art, the beneficial effects of the present invention are:

[0049] 1. This invention proposes a real-time cleaning method for multi-source heterogeneous agricultural data. Through primary cleaning and data format standardization, it reduces reliance on traditional manual cleaning rules and a large amount of manually labeled data, achieves preliminary data error correction, reduces manpower input and maintenance costs, and provides a data foundation for secondary cleaning.

[0050] 2. This invention proposes a real-time cleaning method for multi-source heterogeneous agricultural data. It maps data that has been cleaned and formatted uniformly into a common latent space. By using a customized noise injection and backdiffusion denoising network, the "dirty data" is gradually repaired into a cleaner data expression. The clean data expression is then used as input for repeated iterations. Finally, the optimal data is obtained through evaluation. This method can retain key information while removing noise, significantly improving the data cleaning effect and providing a data foundation for subsequent risk assessment. Attached Figure Description

[0051] Figure 1 Flowchart of the real-time cleaning method for multi-source heterogeneous agricultural data provided by the present invention;

[0052] Figure 2 The secondary cleaning process diagram provided by this invention;

[0053] Figure 3 This is an architecture diagram of the real-time cleaning system for multi-source heterogeneous agricultural data provided by the present invention. Detailed Implementation

[0054] The present application will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present application, but do not limit the present application in any way. It should be noted that those skilled in the art can make several modifications and improvements without departing from the concept of the present application. These all fall within the protection scope of the present application.

[0055] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0056] It should be noted that, unless there is a conflict, the various features in the embodiments of this application can be combined with each other, all of which are within the protection scope of this application. Furthermore, although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than the module division in the device or the order in the flowchart. In addition, the terms "first," "second," and "third" used in this application do not limit the data or execution order, but only distinguish identical or similar items with essentially the same function and effect.

[0057] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The term "and / or" as used in this specification includes any and all combinations of one or more of the associated listed items.

[0058] Example 1

[0059] Please see Figure 1 The present invention provides an embodiment of a real-time cleaning method for multi-source heterogeneous agricultural data, comprising the following specific steps:

[0060] Step S1: Acquire multi-source heterogeneous data, including: satellite remote sensing data, ground flood station observation data, historical disaster data, soil moisture and crop growth data, etc. Clean the acquired multi-source heterogeneous data to obtain cleaned multi-source heterogeneous data.

[0061] In this embodiment, satellite remote sensing data, ground flood station observation data, historical disaster data, soil moisture and crop growth data are acquired through satellite remote sensing, flood stations, physical sensors, etc. These data have the characteristics of being multi-source and heterogeneous, including different data formats such as JSON, XML, CSV, etc., different communication protocols such as HTTP, MQTT, CoAP, etc., and different data sources such as sensor nodes, monitoring platforms, drones, etc.

[0062] The first cleaning of the acquired multi-source heterogeneous data is to perform preliminary screening and filtering of the data to remove obviously abnormal data, such as extreme values ​​caused by sensor failure, such as negative soil moisture values ​​or values ​​exceeding the reasonable range, and records with serious data loss.

[0063] It also includes rule-based cleaning, which involves developing corresponding cleaning rules based on expertise and experience in the agricultural field. These rules include data range verification rules, such as the reasonable range for soil temperature, which is typically between -20℃ and 60℃; data exceeding this range is considered abnormal. Data consistency rules are also included, for example, soil moisture sensor data and rainfall data from weather stations for the same plot are correlated; if rainfall is high but soil moisture does not increase accordingly, there is a data inconsistency problem. Finally, data integrity rules require that key data fields are not missing, such as data collection time and geographical location.

[0064] These cleaning rules are used to clean the acquired multi-source heterogeneous data. For data that violates the value range verification rules, it is handled according to the specific situation. If the outlier is caused by a temporary failure of the acquisition equipment, it is repaired by interpolation using data from adjacent time points. If the data outlier is caused by a long-term failure, it is marked as invalid data. For data inconsistencies, the cause of the inconsistency is found and corrected by comparing and verifying with other relevant data sources. For missing data, interpolation methods such as linear interpolation and polynomial interpolation are used to fill in the gaps.

[0065] Step S2: Convert the multi-source heterogeneous data after one cleaning into a standardized data format, and perform spatiotemporal alignment processing on the data based on a dynamic spatiotemporal grid to obtain the processed multi-source heterogeneous data;

[0066] The specific steps of step S2 are as follows:

[0067] Step S201: Perform feature parsing and field extraction on the multi-source heterogeneous data after one cleaning, analyze the mapping relationship between the data, and establish a data format mapping table;

[0068] In this embodiment, the feature analysis of multi-source heterogeneous data includes, for example: satellite remote sensing data, identified by file extension (.tif / .geotiff) and metadata header information, including features such as the number of bands, spatial resolution (e.g., 10 meters / pixel), projection coordinate system (e.g., WGS84 / GCJ02), and imaging time (timestamp accurate to the second); ground sensor data, determined according to transmission protocol (MQTT / HTTP) and payload format (JSON / CSV), including information such as sensor ID, acquisition frequency (minutes / seconds), and physical quantity unit (soil moisture % / crop height cm); historical disaster data, identified from relational database table structure (field names including "disaster occurrence time" and "latitude and longitude") or Excel file header (e.g., "disaster type" and "affected area"), along with metadata such as data acquisition agency and data version number;

[0069] Field extraction examples: Satellite remote sensing data, spectral data extraction: Read the DN values ​​(digital quantization values) of each band from a remote sensing image parsing library (such as GDAL), convert them into floating-point arrays (dimensions: [number of rows, number of columns, number of bands]); Spatiotemporal metadata parsing: Extract the imaging time (UTC time string), geographic coordinate range (latitude and longitude extremes), and projected coordinate system code (such as EPSG:4326) from the image header file; Ground sensor data, real-time data stream parsing: Extract key-value pairs from the MQTT message payload (JSON format), such as extracting the timestamp and sensor information from {"sensor_id":"A01","time":"2025-04-14T10:00:00Z","humidity":"25.3"}. Location (combined with latitude and longitude information during equipment registration), physical quantity values, offline file parsing: read CSV files line by line, match preset templates by the field names in the first row (e.g., "observation time, water level, flow velocity" corresponds to time, water level value, and flow velocity value); historical disaster data, database query parsing: extract structured data through SQL statements (e.g., SELECT disaster occurrence time, longitude, latitude, disaster level FROM disaster_db), convert it into a two-dimensional table; unstructured text processing: for disaster reports in PDF / Word format, extract key information through natural language processing (NLP) technology (e.g., use regular expressions to match "occurrence time: May 10, 2023" and "location: 118.5°E, 32.0°N");

[0070] A three-level mapping relationship is established based on the data source type (satellite / sensor / historical data). The data format mapping table is formatted as follows: original format → data carrier (file / stream / database) → parsing tool (such as GDAL parsing remote sensing images, Pandas reading CSV) → target field set;

[0071] Step S202: Perform spatiotemporal alignment processing on the data based on the dynamic spatiotemporal grid;

[0072] The specific steps of step S202 are as follows:

[0073] Step S2021: Divide the entire area according to the preset initial standard grid, establish the time axis granularity, and identify sensitive terrain areas and crop planting zones;

[0074] Specifically, the agricultural production area is divided into a standard grid of 500m×500m, the basic time unit is set to 5 minutes, the sensitive terrain area identification automatically adjusts the area with a slope >15° to 100m×100m, and the crop planting zone identification is used to divide the differentiated management area according to the crop type (rice / wheat).

[0075] Step S2022: Establish a UTC standard time axis, calibrate and convert the timestamps in the multi-source heterogeneous data after one cleaning, and mark data of different frequencies with a unified time granularity;

[0076] Different frequency data include high-frequency data and low-frequency data. High-frequency data is such as minute-level data from sensors, and low-frequency data is such as daily data from satellites. The data is marked with a uniform time granularity, for example, with "10 minutes" as the smallest time unit, and equal-interval time series indexes are generated, such as "00:00, 00:10, 00:20...".

[0077] Step S2023: Unify different projected coordinate systems into the WGS84 geographic coordinate system using the seven-parameter transformation method or coordinate transformation library, and add spatial range identifiers to each multi-source heterogeneous data.

[0078] The WGS84 geographic coordinate system uses latitude and longitude. Spatial range identifiers include, for example, satellite data, which uses the latitude and longitude of the pixel center as the reference, with an additional spatial error range, such as ±5 meters; and ground sensors, which use the latitude and longitude of the equipment installation as the reference, with an additional collection coverage radius, such as the average humidity within 5 meters of a soil moisture sensor as the sampling humidity.

[0079] By dynamically adjusting the spatiotemporal reference grid, the problem of integrating "high-frequency point data" and "low-frequency area data" in agricultural IoT is effectively solved, facilitating further data cleaning.

[0080] Step S203: Reconstruct the multi-source heterogeneous data and manage it using dynamic metadata to obtain the processed multi-source heterogeneous data.

[0081] The specific steps of step S203 are as follows:

[0082] Step S2031: Construct a spatiotemporal four-dimensional tensor from the spatiotemporally aligned multi-source heterogeneous data, with dimensions including: longitude, latitude, time, and data source type;

[0083] Step S2032: Create a data traceability archive, record the original data format, acquisition device ID, processing version number, and build a quality labeling system, including integrity labels and consistency labels;

[0084] Among them, the integrity label is a data missing rate of <5%, which is marked as Grade A; the consistency label is a multi-source data difference of <10%, which is marked as Grade B.

[0085] Step S2033: Generate a standard data package, which is the processed multi-source heterogeneous data, including: spatiotemporal four-dimensional tensor, quality label and traceability file.

[0086] In this embodiment, by establishing a standardized data representation system, the data utilization rate is effectively improved, laying a reliable foundation for subsequent secondary data cleaning.

[0087] Step S3: Pre-train the cleaning network. Input the processed multi-source heterogeneous data into the pre-trained cleaning network for iterative cleaning. Randomly select the cleaned multi-source heterogeneous data text during the iteration process and map it to the latent space. Simulate the forward and backward diffusion process of data noise, decode and reconstruct the text, and obtain the final cleaned multi-source heterogeneous data.

[0088] like Figure 2 As shown, the specific steps of step S3 are as follows:

[0089] Step S301: Pre-train the cleaning network, input the processed multi-source heterogeneous data into the pre-trained cleaning network for iterative denoising, randomly select the cleaned multi-source heterogeneous data text during the iteration process and map it to a high-dimensional semantic embedding space, and normalize the embedded vector.

[0090] Specifically, pre-trained cleaning networks such as BERT, GPT, or Transformer aim to transform discrete data text information into continuous representations. The cleaned multi-source heterogeneous data text during the iterative process is the output of the pre-trained cleaning network.

[0091] Step S302: Construct a forward diffusion process by gradually injecting random noise according to a preset noise schedule;

[0092] Specifically, based on the distribution characteristics of the text frontier, a noise sequence is defined, which gradually migrates from a clean latent vector to random noise. This process simulates various interferences or errors encountered by text data during actual collection and transmission, forming a set of "damaged" data with controllable noise levels.

[0093] For example, specific noise injection mechanisms can be designed for different types of data in the potential space, such as using Gaussian noise with mean and variance controlled for image features; for time series and text data, appropriate noise scheduling can be defined according to their distribution characteristics. Specifically, for example, satellite data may have errors due to cloud cover or sensor noise; ground equipment may experience data loss or transmission errors; and manually recorded historical disaster data may have spelling errors or missing information.

[0094] Step S303: Construct a backdiffusion cleaning network to gradually restore clean text vectors from the noisy embedding space;

[0095] The specific steps of step S303 include:

[0096] Step S301: Construct a back-diffusion cleaning network based on the lightweight pre-trained cleaning network;

[0097] Specifically, the backdiffusion cleaning network is a lightweight version of the pre-trained cleaning network. By "slimming down" the redundant parameters or structures of the original cleaning network, the model size is reduced without significantly sacrificing performance.

[0098] Step S302: Set up conditional branches, using the processed multi-source heterogeneous data as conditional inputs to assist the backdiffusion cleaning network in restoring text structure and semantics under different noise levels;

[0099] Step S302 is used to ensure that the original semantics are not lost during the cleaning process and the restoration process;

[0100] Step S303: Use the back diffusion process to gradually restore a clean text representation from the noisy embedding space.

[0101] Specifically, we define mean squared error (MSE) or a reconstruction loss more suited to the characteristics of the text, while introducing semantic consistency constraints.

[0102] Step S304: Use the decoder to map the clean text vector back to the text sequence, and use the text sequence as the input of the pre-trained denoising network for iteration;

[0103] In this embodiment, the decoder is similar to the VAE decoder or adopts a seq2seq structure. After the initial decoding, the text quality is further improved by the post-processing module, such as rule-based syntax correction and language model secondary correction, to ensure that the text structure is rigorous and the language is fluent after data cleaning.

[0104] Step S305: Repeat steps S302-S304 to construct multi-dimensional evaluation metrics, including reconstruction accuracy, grammatical correctness, semantic consistency, etc., evaluate the output of the pre-trained denoising network, and select the data with the best evaluation results as the final cleaned multi-source heterogeneous data.

[0105] exist Figure 2 The text demonstrates a cleaning network and a complete process from vectorization, noise simulation, back diffusion denoising to decoding and reconstruction. The cleaning network generates a result in each iteration, which, together with the result generated by the noise simulation reconstruction, is evaluated and compared to select the optimal data.

[0106] In this embodiment, multi-source heterogeneous data often suffers from noise, missing data, errors, and format mismatches due to differences in acquisition methods, spatiotemporal resolution, and data formats. Traditional text cleaning methods often only focus on local character or word level corrections, making it difficult to simultaneously address low-level errors (such as spelling and word form changes) and high-level semantic losses, ignoring global semantic consistency, resulting in semantic ambiguity and contextual breaks after information recovery.

[0107] In steps S1 and S2, preliminary data cleaning and format unification of multi-source heterogeneous data were completed. Step S3 performs secondary cleaning by simulating text corruption (noise injection) in the latent semantic space and gradually denoising based on a backdiffusion network. It can flexibly adapt to various noise types and levels, and can robustly recover data from minor spelling errors to large-scale information loss. Under different noise levels, it can intelligently allocate recovery resources to ensure that key content is repaired first, which significantly improves the cleaning effect and provides more accurate, robust and complete data support for disaster early warning and risk monitoring.

[0108] Example 2

[0109] Please see Figure 2 Another embodiment of the present invention provides a real-time cleaning system for multi-source heterogeneous agricultural data, comprising: a primary cleaning module, a spatiotemporal processing module, and a secondary cleaning module;

[0110] The first cleaning module is used to acquire multi-source heterogeneous data, perform a first cleaning on the acquired multi-source heterogeneous data, and obtain multi-source heterogeneous data after the first cleaning.

[0111] The spatiotemporal processing module is used to convert multi-source heterogeneous data after one cleaning into a standardized data format, and perform spatiotemporal alignment processing on the data based on a dynamic spatiotemporal grid to obtain processed multi-source heterogeneous data.

[0112] The secondary cleaning module is used to pre-train the cleaning network. The processed multi-source heterogeneous data is input into the pre-trained cleaning network for iterative cleaning. The text of the multi-source heterogeneous data after cleaning during the iteration process is randomly selected and mapped to the latent space to simulate the forward and backward diffusion process of data noise. The text is decoded and reconstructed to obtain the final cleaned multi-source heterogeneous data.

[0113] The spatiotemporal processing module includes: a parsing unit, a spatiotemporal alignment processing unit, and a reconstruction management unit;

[0114] The parsing unit is used to perform feature parsing and field extraction on multi-source heterogeneous data after one cleaning, analyze the mapping relationship between data, and establish a data format mapping table.

[0115] The spatiotemporal alignment processing unit is used to perform alignment processing on multi-source heterogeneous data after one cleaning based on time and space.

[0116] The reconstruction management unit is used to construct a spatiotemporal four-dimensional tensor and a data traceability archive, and to generate a standard data package.

[0117] The secondary cleaning module includes: a vectorization processing unit, a forward diffusion unit, a reverse diffusion cleaning unit, a decoding unit, and a cleaning evaluation unit;

[0118] The vectorization processing unit is used to input the processed multi-source heterogeneous data into a pre-trained cleaning network for iterative denoising, randomly select the cleaned multi-source heterogeneous data text during the iteration process, map it to a high-dimensional semantic embedding space, and normalize the embedded vector.

[0119] The forward diffusion unit is used to construct the forward diffusion process and gradually inject random noise according to a preset noise schedule.

[0120] The backdiffusion cleaning unit is used to construct a backdiffusion cleaning network to gradually restore clean text vectors from a noisy embedding space.

[0121] The decoding unit is used to map clean text vectors back to text sequences using a decoder, and to iterate the text sequences as input to a pre-trained denoising network.

[0122] The cleaning and evaluation unit is used to evaluate the output of the pre-trained denoising network and select the data with the best evaluation result as the final cleaned multi-source heterogeneous data.

[0123] In addition, the parts of the technical solutions provided in the embodiments of this application that are consistent with the implementation principles of the corresponding technical solutions in the prior art have not been described in detail, so as to avoid excessive elaboration.

[0124] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the invention. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for real-time cleaning of multi-source heterogeneous agricultural data, characterized in that, include: Acquire multi-source heterogeneous data, perform cleaning on the acquired multi-source heterogeneous data, and obtain cleaned multi-source heterogeneous data. After cleaning, the multi-source heterogeneous data is converted into a standardized data format, and spatiotemporal alignment is performed on the data based on a dynamic spatiotemporal grid to obtain the processed multi-source heterogeneous data. A pre-trained cleaning network is used to input the processed multi-source heterogeneous data into the pre-trained cleaning network for iterative cleaning. The text of the multi-source heterogeneous data cleaned during the iteration process is randomly selected and mapped to the latent space to simulate the forward and backward diffusion of data noise. The text is then decoded and reconstructed to obtain the final cleaned multi-source heterogeneous data. The pre-trained cleaning network inputs the processed multi-source heterogeneous data into the pre-trained cleaning network for iterative cleaning. During the iteration process, it randomly selects the cleaned multi-source heterogeneous data text and maps it to the latent space, simulating the forward and backward diffusion processes of data noise. The text is then decoded and reconstructed, including: A pre-trained cleaning network is used to input the processed multi-source heterogeneous data into the pre-trained cleaning network for iterative denoising. The text of the multi-source heterogeneous data after cleaning during the iteration process is randomly selected and mapped to a semantic embedding space. The embedded vector is then normalized. Construct a forward diffusion process and gradually inject random noise according to a preset noise schedule; Construct a backdiffusion cleaning network to gradually restore clean text vectors from a noisy embedding space; The decoder is used to map clean text vectors back to text sequences, and the text sequences are used as input to a pre-trained denoising network for iterative processing. The process involves repeated random noise injection, restoration, and mapping iterations to construct multi-dimensional evaluation metrics, including reconstruction accuracy, syntactic correctness, and semantic consistency. The output of the pre-trained denoising network is then evaluated, and the data with the best evaluation results is selected as the final cleaned multi-source heterogeneous data. The construction of the backdiffusion cleaning network, which progressively restores clean text vectors from a noisy embedding space, includes: Based on a lightweight pre-trained cleaning network, a back-diffusion cleaning network is constructed. By setting conditional branches and using the processed multi-source heterogeneous data as conditional inputs, the backdiffusion cleaning network is assisted in restoring the text structure and semantics under different noise levels. A clean text representation is gradually restored from a noisy embedding space using a backdiffusion process.

2. The real-time cleaning method for multi-source heterogeneous agricultural data as described in claim 1, characterized in that, The process of converting multi-source heterogeneous data after a single cleaning into a standardized data format and performing spatiotemporal alignment processing on the data based on a dynamic spatiotemporal grid includes: Feature parsing and field extraction are performed on multi-source heterogeneous data after one cleaning process, the mapping relationship between data is analyzed, and a data format mapping table is established. Data is spatiotemporally aligned based on a dynamic spatiotemporal grid. Multi-source heterogeneous data is reconstructed and managed using dynamic metadata to obtain processed multi-source heterogeneous data.

3. The real-time cleaning method for multi-source heterogeneous agricultural data as described in claim 2, characterized in that, The spatiotemporal alignment processing of data based on a dynamic spatiotemporal grid includes: The entire region is divided according to a preset initial standard grid, a time axis granularity is established, and sensitive terrain areas and crop planting zones are identified. Establish a UTC standard timeline, calibrate and convert timestamps in multi-source heterogeneous data after one cleaning, and mark data of different frequencies with a unified time granularity. Different projected coordinate systems are unified into the WGS84 geographic coordinate system using the seven-parameter transformation method or coordinate transformation library, and spatial range identifiers are added to each multi-source heterogeneous data.

4. The real-time cleaning method for multi-source heterogeneous agricultural data as described in claim 2, characterized in that, The process of reconstructing multi-source heterogeneous data and managing it using dynamic metadata to obtain processed multi-source heterogeneous data includes: The spatiotemporally aligned multi-source heterogeneous data is constructed into a spatiotemporal four-dimensional tensor, with dimensions including: longitude, latitude, time, and data source type. Create data traceability archives to record the original data format, acquisition device ID, and processing version number; and build a quality labeling system, including integrity labels and consistency labels. Generate standard data packets, which are processed multi-source heterogeneous data, including: spatiotemporal four-dimensional tensors, quality labels, and traceability archives.

5. The real-time cleaning method for multi-source heterogeneous agricultural data as described in claim 1, characterized in that, The multi-source heterogeneous data includes: satellite remote sensing data, ground flood station observation data, historical disaster data, soil moisture and crop growth data.

6. A real-time cleaning system for multi-source heterogeneous agricultural data, used to implement the real-time cleaning method for multi-source heterogeneous agricultural data as described in any one of claims 1-5, characterized in that, include: The system consists of a primary cleaning module, a spatiotemporal processing module, and a secondary cleaning module. The first cleaning module is used to acquire multi-source heterogeneous data, perform a first cleaning on the acquired multi-source heterogeneous data, and obtain multi-source heterogeneous data after the first cleaning. The spatiotemporal processing module is used to convert multi-source heterogeneous data after one cleaning into a standardized data format, and perform spatiotemporal alignment processing on the data based on a dynamic spatiotemporal grid to obtain processed multi-source heterogeneous data. The secondary cleaning module is used to pre-train the cleaning network. The processed multi-source heterogeneous data is input into the pre-trained cleaning network for iterative cleaning. The text of the multi-source heterogeneous data after cleaning during the iteration process is randomly selected and mapped to the latent space to simulate the forward and backward diffusion process of data noise. The text is decoded and reconstructed to obtain the final cleaned multi-source heterogeneous data.

7. The real-time cleaning system for multi-source heterogeneous agricultural data as described in claim 6, characterized in that, The spatiotemporal processing module includes: a parsing unit, a spatiotemporal alignment processing unit, and a reconstruction management unit; The parsing unit is used to perform feature parsing and field extraction on multi-source heterogeneous data after one cleaning, analyze the mapping relationship between data, and establish a data format mapping table. The spatiotemporal alignment processing unit is used to perform alignment processing on multi-source heterogeneous data after one cleaning based on time and space. The reconstruction management unit is used to construct a spatiotemporal four-dimensional tensor and a data traceability archive, and to generate a standard data package.

8. The real-time cleaning system for multi-source heterogeneous agricultural data as described in claim 7, characterized in that, The secondary cleaning module includes: a vectorization processing unit, a forward diffusion unit, a reverse diffusion cleaning unit, a decoding unit, and a cleaning evaluation unit; The vectorization processing unit is used to input the processed multi-source heterogeneous data into a pre-trained cleaning network for iterative denoising, randomly select the cleaned multi-source heterogeneous data text during the iteration process, map it to a semantic embedding space, and normalize the embedded vector. The forward diffusion unit is used to construct the forward diffusion process and gradually inject random noise according to a preset noise schedule. The backdiffusion cleaning unit is used to construct a backdiffusion cleaning network to gradually restore clean text vectors from a noisy embedding space. The decoding unit is used to map clean text vectors back to text sequences using a decoder, and to iterate the text sequences as input to a pre-trained denoising network. The cleaning and evaluation unit is used to evaluate the output of the pre-trained denoising network and select the data with the best evaluation result as the final cleaned multi-source heterogeneous data.

Citation Information

Patent Citations

  • Fire-fighting information intelligent acquisition and management method based on Internet of Things

    CN117828280A

  • Land space planning environment influence monitoring method and device

    CN119848703A

  • Fusion system and method for multi-source heterogeneous agricultural data elements

    CN120105315A