Method and system for constructing agricultural variety adaptation dataset based on multi-source heterogeneous data

By standardizing and indexing multi-source heterogeneous agricultural data, the problems of fragmented data standardization rules and missing cross-source associations were solved, generating a standardized agricultural variety matching dataset and realizing the orderly collection and high-quality fusion of data.

CN122633771APending Publication Date: 2026-08-25SINOCHEM INFORMATION TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610983952.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-03
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

In existing technologies, the processing of agricultural variety-related data suffers from fragmented data standardization rules, a lack of unified index support for cross-source association, crude fusion logic, and a lack of data quality control system, making it difficult to generate standardized and usable agricultural variety-adapted datasets.

Method used

By acquiring multi-source heterogeneous agricultural data, a differentiated processing strategy is adopted for standardization. Administrative division code index and crop type index are configured, and data association and fusion are carried out based on the dual index system. Consistency verification, integrity verification and accuracy verification are implemented, and finally standardized asset encapsulation is performed to generate standardized data access interfaces.

Benefits of technology

It has achieved standardized transformation and unification of multi-source heterogeneous agricultural data, ensuring the logic, standardization and uniformity of data fusion, improving data quality and generating usable agricultural variety adaptation datasets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122633771A_ABST
    Figure CN122633771A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of intelligent agriculture, and discloses an agricultural variety adaptation dataset construction method and system based on multi-source heterogeneous data, which comprises the following steps: standardizing first multi-source heterogeneous agricultural data by using a differentiated processing strategy to obtain second multi-source heterogeneous agricultural data, which is associated with a first index and a second index respectively to obtain third multi-source heterogeneous agricultural data; generating a first fusion dataset based on a first fusion strategy and the third multi-source heterogeneous agricultural data; sequentially performing consistency checking, integrity checking and accuracy checking to obtain a second fusion dataset; and performing standardized asset packaging to generate a standardized data access interface and an agricultural variety adaptation dataset. The method solves the problems of the related technologies, such as scattered data standardization rules, lack of unified index support for cross-source association, extensive fusion logic, and missing data quality control system, and can output a standard and usable agricultural variety adaptation dataset.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart agriculture technology, specifically to a method and system for constructing an agricultural variety adaptation dataset based on multi-source heterogeneous data. Background Technology

[0002] The regional promotion of improved varieties in modern agriculture is a core link in ensuring stable and efficient grain production in my country. Whether crop varieties can be accurately matched with the soil, climate, and pest and disease endowments of the planting area directly determines the planting benefits and regional agricultural productivity. Currently, the digital transformation of the seed industry in China is progressing steadily. Relevant departments such as agriculture and rural affairs, meteorology, and natural resources have implemented special data construction projects, gradually accumulating massive amounts of basic data in four categories: soil sampling, meteorological time series, pest and disease records, and variety approval. Soil sampling data is collected annually based on the national soil survey; meteorological time series data is generated from continuous monitoring by meteorological stations at all levels; unstructured pest and disease data comes from field monitoring records of grassroots plant protection departments and industry research literature; and seed approval data is archived by the National Agricultural Technology Extension Service Center based on the results of variety trials. These four types of data are managed and maintained by different agencies, and their data sources, storage formats, and field specifications are inconsistent, constituting typical multi-source heterogeneous agricultural data. Building standardized variety adaptation datasets based on the above data has become an essential requirement for variety introduction zoning and new variety approval auxiliary analysis. The existing first-generation method for organizing variety adaptation data is a manual ledger archiving model, which was also the common solution for grassroots agricultural technology extension agencies in the early days. This model can only collect small-scale, scattered sampling data and can only support small-batch variety trials in counties, lacking the ability to build large-scale datasets. With the popularization of digital office work, the industry has evolved to the second-generation technical solution of independent databases built by different departments. Each competent unit builds its own dedicated database, the natural resources department stores soil sampling data separately, meteorological agencies independently operate meteorological time-series databases, plant protection units archive unstructured documents on pests and diseases, and agricultural technology departments keep seed approval ledgers separately. Each database can complete simple standardization processing of single-type data and can assign administrative division codes to the data in its own database, but the entire industry lacks a unified administrative division index and standardized crop type index based on the national standard GB / T2260. There are significant differences in the compilation standards of administrative division codes between different databases. Some use Chinese place names for labeling, while others use custom short codes. Crop classification description rules are inconsistent, and standardized association matching cannot be carried out between databases. Various types of data have long been in a state of data silos, making it difficult to achieve cross-source data linkage and collection. Currently, the mainstream approach in the industry is to use a third-generation simplified data fusion scheme that uses literal field names for concatenation. Some agricultural big data platforms are attempting to break down data silos by matching the literal names of database field names to merge multiple types of data. However, this type of fusion method does not pre-build a dual-keyword fusion strategy based on regional indexes and crop indexes. It relies solely on mechanically piecing together data using field names. The resulting data entries generally suffer from problems such as mixing data from different crops in the same region and mismatches between approved varieties and planting environments, failing to generate standardized and unified integrated data entries.

[0003] In summary, the data processing technologies for related agricultural varieties suffer from multiple technical shortcomings, including fragmented data standardization rules, lack of unified index support for cross-source association, crude fusion logic, and absence of a data quality control system, making it difficult to produce standardized and usable agricultural variety-adapted datasets. Summary of the Invention

[0004] In view of this, the present invention provides a method and system for constructing an agricultural variety adaptation dataset based on multi-source heterogeneous data, in order to solve the problem that related agricultural variety data processing technologies have multiple technical shortcomings, such as scattered data standardization rules, lack of unified index support for cross-source association, crude fusion logic, and lack of data quality control system, making it difficult to produce standardized and usable agricultural variety adaptation datasets.

[0005] In a first aspect, the present invention provides a method for constructing an agricultural variety adaptation dataset based on multi-source heterogeneous data. The method includes: acquiring first multi-source heterogeneous agricultural data, which includes soil sampling data, meteorological time-series data, unstructured pest and disease data, and basic seed approval data; standardizing the first multi-source heterogeneous agricultural data using a differentiated processing strategy to obtain second multi-source heterogeneous agricultural data, which includes standardized soil sampling data, standardized meteorological time-series data, standardized unstructured pest and disease data, and standardized basic seed approval data; acquiring a first index and a second index, and associating the second multi-source heterogeneous agricultural data with the first and second indexes respectively to obtain third multi-source heterogeneous agricultural data, where the first index is an administrative division code index and the second index is a crop type index; acquiring a first fusion strategy, and generating a first fusion dataset based on the first fusion strategy and the third multi-source heterogeneous agricultural data, where the first fusion strategy is pre-constructed based on the first and second indexes; sequentially performing consistency verification, integrity verification, and accuracy verification on the first fusion dataset to obtain a second fusion dataset; and performing standardized asset encapsulation on the second fusion dataset to generate a standardized data access interface and an agricultural variety adaptation dataset.

[0006] The method for constructing an agricultural variety adaptation dataset based on multi-source heterogeneous data provided in this embodiment firstly covers the core data dimensions involved in the crop growth adaptation process, such as soil and water environment, climate conditions, pest and disease impacts, and variety approval qualifications, by collecting soil sampling data, meteorological time series data, unstructured pest and disease data, and seed basic approval data. This fully covers the basic data system required for agricultural variety adaptation assessment. The multi-dimensional data acquisition model fully meets the underlying data needs of agricultural variety adaptation scenarios, avoiding the information gaps caused by single data dimensions. Based on multi-dimensional raw data, it provides complete and business-scenario-appropriate raw data support for subsequent data standardization processing, data association and fusion, and other full-process technical operations. Secondly, it adopts differentiated processing strategies to carry out targeted standardization processing of multi-source heterogeneous agricultural data. Based on the structural characteristics, data forms, and information carrier differences of different types of data, it configures corresponding data regularization logic. For structured soil and water data, meteorological data, and seed data, it performs parameter extraction and regularization processing, and for unstructured pest and disease data, it performs annotation and standardization sorting. It adapts to the own attribute characteristics of various heterogeneous data, solves the technical problems of inconsistent formats, disordered parameters, and mismatched information regularization logic of different data sources, realizes the standardized conversion of various raw heterogeneous data, and unifies the basic form of multi-source data. Then, by configuring two standardized index systems—administrative division code index and crop type index—two-way dimensional association and binding operations are carried out on the standardized multi-source agricultural data. A unified data association benchmark is built based on these two dedicated indexes, assigning unified spatial attribution and crop classification identifiers to all types of agricultural data, and establishing an underlying association logic that enables interoperability and matching between heterogeneous data across categories. The standardized index configuration mode breaks down the association barriers formed by the independent storage and organization of various types of data, providing a technical foundation for linked matching of scattered and independent multi-source data, achieving orderly aggregation and dimensional unification of multi-source heterogeneous data. Subsequently, by pre-constructing a dedicated first fusion strategy based on the dual index system, the fusion process of multi-source data is coordinated based on standardized dual-keyword matching logic, abandoning the disorderly splicing processing mode. Data grouping, adaptive data filtering, and variety information matching operations are completed based on the index dimension. Environmental data aggregation and variety data pairing are strictly completed according to the dual logic of spatial and crop dimensions, standardizing the overall execution logic of multi-source data fusion. An orderly data fusion model enables the orderly transformation of fragmented single-category data into integrated data, ensuring the logic, standardization, and consistency of the data fusion process, and avoiding logical confusion and information mismatch problems in the data fusion process.Furthermore, through progressive quality control operations involving consistency, integrity, and accuracy checks on the initial fusion dataset, data quality is optimized layer by layer from multiple technical levels, including field logical relationships, completeness of core parameters, and overall data authenticity. This process systematically addresses and corrects internal logical conflicts, fills in missing core parameters, and standardizes overall data accuracy. A systematic data quality control logic is built based on multi-level verification methods, and this progressive verification model comprehensively optimizes the internal data structure and content quality of the dataset. Finally, standardized asset encapsulation is performed on the fusion dataset after quality verification. This unifies the dataset's field definitions, storage formats, and data specifications, establishes a multi-dimensional retrieval system, and generates standardized data access interfaces. The standardized asset encapsulation model enables the compliant transformation of fragmented data resources into standardized, structured digital assets, establishes a unified data retrieval and calling mechanism, standardizes the output format and usage of the dataset, and achieves standardized storage, convenient retrieval, and stable reuse of agricultural variety adaptation datasets based on a standardized interface system. By implementing this solution, we have solved the problem that agricultural variety-related data processing technologies suffer from multiple technical shortcomings, such as fragmented data standardization rules, lack of unified index support for cross-source association, crude fusion logic, and lack of a data quality control system, making it difficult to produce standardized and usable agricultural variety-adapted datasets.

[0007] In one optional implementation, acquiring first multi-source heterogeneous agricultural data includes: acquiring soil sampling data based on the national soil survey database; acquiring meteorological time-series data based on the public interface of the National Meteorological Center and third-party meteorological service platforms; acquiring unstructured pest and disease data based on pest and disease forecasting data from the Ministry of Agriculture and Rural Affairs and publicly available online information; and acquiring basic seed approval data based on variety approval announcements from the National Agricultural Technology Extension Service Center.

[0008] In one optional implementation, a differentiated processing strategy is used to standardize the first multi-source heterogeneous agricultural data to obtain the second multi-source heterogeneous agricultural data. This includes: extracting a first parameter set from the soil sampling data, the first parameter set including pH value, organic matter content, total nitrogen content, available phosphorus content, and available potassium content; mapping the discrete sampling point data corresponding to the first parameter set to the corresponding geographic grid; assigning administrative division codes to each geographic grid; using the IQR outlier detection algorithm to remove outlier sampling values ​​in the first parameter set; and using spatial interpolation to complete the missing values ​​of the first parameter set in each geographic grid to obtain standardized soil sampling data.

[0009] In one optional implementation, a differentiated processing strategy is used to standardize the first multi-source heterogeneous agricultural data to obtain the second multi-source heterogeneous agricultural data. This includes: extracting a second set of parameters from the meteorological time-series data, the second set of parameters including effective accumulated temperature, active accumulated temperature, effective precipitation, and active precipitation; assigning administrative division codes to the second set of parameters based on meteorological station data; and classifying and integrating the second set of parameters according to the administrative division codes to obtain standardized meteorological time-series data.

[0010] In one optional implementation, a differentiated processing strategy is used to standardize the first multi-source heterogeneous agricultural data to obtain the second multi-source heterogeneous agricultural data. This includes: acquiring a pre-constructed standardized knowledge graph of pests and diseases, which includes disease name, pest name, pathogen, source of pests, stage of damage, and control threshold; assigning administrative division codes to the unstructured data of pests and diseases; and labeling the unstructured data of pests and diseases based on the administrative division codes and the standardized knowledge graph of pests and diseases to obtain standardized unstructured data of pests and diseases.

[0011] In one optional implementation, a differentiated processing strategy is used to standardize the first multi-source heterogeneous agricultural data to obtain the second multi-source heterogeneous agricultural data. This includes: extracting a third parameter set from the seed basic approval data, the third parameter set including variety name, approval number, approval area, suitable planting season, and characteristic features; assigning administrative division codes to the approval areas in the third parameter set; identifying and removing duplicate approval information in the third parameter set; and classifying and archiving the third parameter set according to crop type to obtain standardized seed basic approval data.

[0012] In one optional implementation, a first index and a second index are obtained, and the second multi-source heterogeneous agricultural data are associated with the first index and the second index respectively to obtain the third multi-source heterogeneous agricultural data. This includes: obtaining the first index and the second index based on a public platform; performing administrative division-level association matching between standardized soil sampling data, standardized meteorological time-series data, standardized pest and disease unstructured data, and standardized seed basic approval data and the first index respectively; and performing crop type-level association matching between the standardized soil sampling data, standardized meteorological time-series data, standardized pest and disease unstructured data, and standardized seed basic approval data that have undergone administrative division-level association matching, and the second index respectively to generate the third multi-source heterogeneous agricultural data.

[0013] In one optional implementation, a first fusion strategy is obtained, and a first fusion dataset is generated based on the first fusion strategy and the third multi-source heterogeneous agricultural data. This includes: obtaining the first fusion strategy, which is a dual-keyword matching fusion rule; based on the dual-keyword matching fusion rule, using a first index and administrative division codes, grouping and fusion within groups the standardized soil sampling data, standardized meteorological time-series data, and standardized pest and disease unstructured data in the third multi-source heterogeneous agricultural data to obtain several groups of first grouped data; based on the dual-keyword matching fusion rule, using a second index, matching crop types in each group of first grouped data to obtain environmental adaptation data corresponding to each crop type; performing item-by-item matching based on the environmental adaptation data corresponding to each crop type and seed basic approval data to obtain several integrated data entries; and constructing the first fusion dataset based on these integrated data entries.

[0014] In one optional implementation, consistency verification, integrity verification, and accuracy verification are sequentially performed on the first fused dataset to obtain the second fused dataset. This includes: verifying the logical matching relationship of the administrative division code, geographic coordinates, and crop type fields in the first fused dataset, and removing data entries with logical field conflicts and coordinate errors to obtain the first fused dataset after consistency verification; using the mean of data from similar regions and model inference to complete the missing core parameter fields within the data entries in the first fused dataset after consistency verification, the core parameter fields include a first parameter set, a second parameter set, pest and disease risk parameters, and a third parameter set to obtain the first fused dataset after integrity verification; and performing accuracy verification on the first fused dataset after integrity verification to generate the second fused dataset. The accuracy verification includes sampling comparison verification with authoritative agricultural databases and manual sampling verification by agricultural experts.

[0015] Secondly, this invention provides a system for constructing an agricultural variety adaptation dataset based on multi-source heterogeneous data. The system includes: an acquisition module for acquiring first multi-source heterogeneous agricultural data, which includes soil sampling data, meteorological time-series data, unstructured pest and disease data, and basic seed approval data; a processing module for standardizing the first multi-source heterogeneous agricultural data using a differentiated processing strategy to obtain second multi-source heterogeneous agricultural data, which includes standardized soil sampling data, standardized meteorological time-series data, standardized unstructured pest and disease data, and standardized basic seed approval data; and an association module for acquiring a first index and a second index. The system consists of three modules: a first index and a second index; a third index; a fusion module; a first fusion strategy; a second fusion dataset; a verification module; and a generation module. The second fusion dataset is then used to perform consistency, integrity, and accuracy checks on the first fusion dataset sequentially to obtain the second fusion dataset. Attached Figure Description

[0016] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0017] Figure 1 This is a flowchart illustrating a specific example of a method for constructing an agricultural variety adaptation dataset based on multi-source heterogeneous data according to an embodiment of the present invention. Figure 2 This is a schematic diagram of a specific example of an agricultural variety adaptation dataset construction system based on multi-source heterogeneous data according to an embodiment of the present invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] The regional promotion of improved varieties in modern agriculture is a core link in ensuring stable and efficient grain production in my country. Whether crop varieties can be accurately matched with the soil, climate, and pest and disease endowments of the planting area directly determines the planting benefits and regional agricultural productivity. Currently, the digital transformation of the seed industry in China is progressing steadily. Relevant departments such as agriculture and rural affairs, meteorology, and natural resources have implemented special data construction projects, gradually accumulating massive amounts of basic data in four categories: soil sampling, meteorological time series, pest and disease records, and variety approval. Soil sampling data is collected annually based on the national soil survey; meteorological time series data is generated from continuous monitoring by meteorological stations at all levels; unstructured pest and disease data comes from field monitoring records of grassroots plant protection departments and industry research literature; and seed approval data is archived by the National Agricultural Technology Extension Service Center based on the results of variety trials. These four types of data are managed and maintained by different agencies, and their data sources, storage formats, and field specifications are inconsistent, constituting typical multi-source heterogeneous agricultural data. Building standardized variety adaptation datasets based on the above data has become an essential requirement for variety introduction zoning and new variety approval auxiliary analysis. The existing first-generation method for organizing variety adaptation data is a manual ledger archiving model, which was also the common solution for grassroots agricultural technology extension agencies in the early days. This model can only collect small-scale, scattered sampling data and can only support small-batch variety trials in counties, lacking the ability to build large-scale datasets. With the popularization of digital office work, the industry has evolved to the second-generation technical solution of independent databases built by different departments. Each competent unit builds its own dedicated database, the natural resources department stores soil sampling data separately, meteorological agencies independently operate meteorological time-series databases, plant protection units archive unstructured documents on pests and diseases, and agricultural technology departments keep seed approval ledgers separately. Each database can complete simple standardization processing of single-type data and can assign administrative division codes to the data in its own database, but the entire industry lacks a unified administrative division index and standardized crop type index based on the national standard GB / T2260. There are significant differences in the compilation standards of administrative division codes between different databases. Some use Chinese place names for labeling, while others use custom short codes. Crop classification description rules are inconsistent, and standardized association matching cannot be carried out between databases. Various types of data have long been in a state of data silos, making it difficult to achieve cross-source data linkage and collection. Currently, the mainstream approach in the industry is to use a third-generation simplified data fusion scheme that uses literal field names for concatenation. Some agricultural big data platforms are attempting to break down data silos by matching the literal names of database field names to merge multiple types of data. However, this type of fusion method does not pre-build a dual-keyword fusion strategy based on regional indexes and crop indexes. It relies solely on mechanically piecing together data using field names. The resulting data entries generally suffer from problems such as mixing data from different crops in the same region and mismatches between approved varieties and planting environments, failing to generate standardized and unified integrated data entries.

[0020] In summary, the data processing technologies for related agricultural varieties suffer from multiple technical shortcomings, including fragmented data standardization rules, lack of unified index support for cross-source association, crude fusion logic, and absence of a data quality control system, making it difficult to produce standardized and usable agricultural variety-adapted datasets.

[0021] To address the technical problems mentioned in the background section, this application provides a method for constructing an agricultural variety adaptation dataset based on multi-source heterogeneous data. See details below. Figure 1 As shown, the method includes: Step S101: Obtain the first multi-source heterogeneous agricultural data, which includes soil sampling data, meteorological time series data, unstructured data on pests and diseases, and seed basic approval data.

[0022] Specifically, step S101 includes: Step a1: Obtain soil sampling data based on the national soil survey database.

[0023] Furthermore, based on the National Soil Survey Database, a full retrieval and content screening of soil sampling data was initiated. This database is a specialized public database established by natural resources and agricultural authorities to coordinate soil resource surveys. It contains monitoring data from field sampling conducted in batches across the country. The screening process identified all types of original soil records needed for subsequent standardization, removing redundant supplementary materials such as project initiation documents and field survey logs irrelevant to soil parameters. After screening, the original records were exported in batches and then uniformly categorized and stored to form complete soil sampling data. During data collection in the rice-growing areas of the middle and lower reaches of the Yangtze River, original records generated from field sampling in various districts and counties along the river were retrieved from the corresponding database sections. Soil physicochemical records were screened and retained, while irrelevant documents such as project applications and field notes were removed. Based on the screened and retained content, soil sampling data for designated areas was collected.

[0024] Step a2: Obtain meteorological time-series data based on the public interface of the National Meteorological Center and third-party meteorological service platforms.

[0025] Furthermore, the National Meteorological Center's public interface serves as a data retrieval channel for authoritative domestic meteorological authorities, while third-party meteorological service platforms provide data services to market-oriented institutions with compliant meteorological data collection qualifications. Climate monitoring data uploaded daily from ground meteorological monitoring stations at all levels is retrieved periodically from the National Meteorological Center's public interface. Detailed regional small-scale climate observation data is supplemented from the third-party meteorological service platforms. The monitoring data obtained from both channels is uniformly summarized and integrated. After removing invalid and missing records, the integrated data is aggregated to form meteorological time-series data. When collecting data for the main winter wheat producing areas in North China, annual monitoring data from regional national meteorological stations is retrieved from the official interface. Simultaneously, compliant and formal meteorological service providers supplement detailed daily climate records at the county level, eliminating blank and invalid records. The two types of valid data are merged and organized to generate corresponding regional meteorological time-series data.

[0026] Step a3: Obtain unstructured pest and disease data based on pest and disease monitoring data from the Ministry of Agriculture and Rural Affairs and publicly available online information.

[0027] Furthermore, the pest and disease monitoring data from the Ministry of Agriculture and Rural Affairs consists of official archived data compiled from regular reports on field pest and disease occurrences by plant protection agencies across the country. Publicly available online data includes textual materials such as field pest and disease survey reports published on agricultural science popularization platforms and local agricultural technology journals. Standardized monitoring reports were downloaded in batches from the official monitoring archives, and pest and disease survey content published on legitimate agricultural websites was filtered through compliant web scraping methods. All textual materials from both sources were collected and summarized without structural field splitting, preserving the original text format to form unstructured pest and disease data. During the data collection phase for rice and cash crops in South China, monthly pest and disease monitoring archives from the Plant Protection Station were retrieved, and field pest and disease survey notes published on local agricultural science and education platforms were simultaneously scraped. Without splitting text fields, all original text content was summarized to form unstructured pest and disease data for the corresponding regions.

[0028] Step a4: Obtain basic seed approval data based on the variety approval announcements from the National Agricultural Technology Extension Service Center.

[0029] Furthermore, the National Agricultural Technology Extension Service Center is the official management unit that coordinates the domestic crop variety testing and approval business. Variety approval announcements are publicized documents confirming the compliance of varieties after regional introduction trials and verification. The original texts of official announcements for different crop categories are retrieved batch by batch. Text relevant to variety registration information is extracted from the original announcements, while irrelevant sections such as policy explanations and public notices are removed. All extracted registration information is then centrally collected and organized to form basic seed approval data. When collecting data on summer maize varieties in the Huang-Huai-Hai Plain, historical maize variety approval announcements are retrieved in batches. Variety names, approval regions, and variety characteristics are extracted from these documents. Additional management clauses and public notices are removed. Based on the selected content, basic seed approval data for the corresponding variety category is compiled.

[0030] Step S102: Standardize the first multi-source heterogeneous agricultural data using a differentiated processing strategy to obtain the second multi-source heterogeneous agricultural data. The second multi-source heterogeneous agricultural data includes standardized soil sampling data, standardized meteorological time-series data, standardized unstructured pest and disease data, and standardized seed basic approval data.

[0031] Specifically, step S102 includes: Step b1: Extract the first set of parameters from the soil sampling data. The first set of parameters includes pH value, organic matter content, total nitrogen content, available phosphorus content, and available potassium content.

[0032] Furthermore, soil sampling data encompasses various supplementary materials, including field sampling records and auxiliary notes accompanying the sampling operation. The first parameter set refers to five core physicochemical indicators used to assess the suitability of soil properties in a plot. These five indicators correspond to the plot's pH level, soil organic matter retention level, soil total element retention level, available phosphorus enrichment level, and available potassium enrichment level, respectively. All original records were meticulously reviewed, and redundant information irrelevant to the five physicochemical indicators, such as information on sampling personnel, field sampling itineraries, and instrument maintenance records, was removed from the disorganized raw data. Fields corresponding to each of the five indicators were individually extracted, and all selected fields were compiled to form a complete first parameter set. When organizing soil data in the main winter wheat producing area of ​​the Huang-Huai-Hai Plain, the entire area's field sampling archives were reviewed line by line, irrelevant notes were removed, and the five soil parameter information was extracted to construct the corresponding first parameter set for each area.

[0033] Step b2: Map the discrete sampling point data corresponding to the first parameter set to the corresponding geographic grid.

[0034] Furthermore, the geographic grid refers to contiguous spatial units formed by uniformly dividing the land area according to geographic spatial rules. Discrete sampling point data represents the accompanying first parameter set information collected from random points in the field, and each discrete sampling point carries its own unique geographic location identifier. The spatial location information corresponding to each discrete sampling point bound to the first parameter set is retrieved, and compared with the pre-defined boundaries of the global geographic grid, the spatial range of each individual discrete sampling point is matched one by one, and the first parameter set content carried by the corresponding point is assigned to the corresponding geographic grid to which the point belongs. The spatial matching and classification of all scattered points within the area is completed one by one, without omitting any valid discrete sampling record. During the processing of soil data for rice-growing areas in the middle and lower reaches of the Yangtze River, based on the geographical location of the points, the information of the scattered sampling points along the river fields is assigned one by one to the corresponding geographic grid.

[0035] Step b3: Assign administrative division codes to each geographic grid.

[0036] Furthermore, administrative division codes refer to the unique identifiers based on the unified national regional division standards, with different administrative jurisdictions corresponding to unique codes. The process involves retrieving the baseline resources for administrative division codes across the entire region, retrieving the regional affiliation information for each geographic grid that has already incorporated sampling parameters, matching the appropriate administrative division codes against the regional affiliation information, and attaching the matched administrative division codes to all geographic grids. All geographic grids within the same administrative jurisdiction are uniformly configured with the same administrative division codes, and the geographic grids after coding configuration are synchronously bound to their own first set of incorporated parameters. When processing grid data from the hilly dryland grain planting area in Central China, the administrative division codes corresponding to the relevant districts and counties are configured for all farmland geographic grids within the area according to the grid regional affiliation results.

[0037] Step b4: Use the IQR outlier detection algorithm to remove outlier sampled values ​​from the first parameter set.

[0038] Furthermore, the IQR outlier detection algorithm refers to a data screening method based on quartile interval rules to identify records that deviate from the normal data distribution range. It aggregates all first-parameter set data within a single geographic grid, performs numerical distribution statistics on all records of five soil parameters within the grid, defines a reasonable data fluctuation range based on the statistically obtained interval boundaries, compares the matching status of each sampled parameter with the defined range, and determines sampled content that falls outside the reasonable fluctuation range as anomaly. All identified anomaly samples are directly removed from the first-parameter set, retaining only compliant parameter records within the reasonable range. When processing soil parameters in the Northeast black soil planting area, the corresponding algorithm is used to screen for distorted measured data caused by field environmental interference, removing all anomaly records and retaining the valid soil parameters for the area.

[0039] Step b5 involves using spatial interpolation to fill in the missing values ​​of the first parameter set within each geographic grid, thereby obtaining standardized soil sampling data.

[0040] Furthermore, spatial interpolation refers to the data supplementation method that, based on the existing parameter content of surrounding valid spatial points, extrapolates and fills in the missing parameters of adjacent plots. It involves retrieving the first parameter set content retained within each geographic grid, screening for records with missing parameters, and identifying all grids with missing content and their corresponding missing parameter categories. Valid first parameter set data retained within adjacent geographic grids are retrieved, and spatial interpolation is used to fill in the missing parameter information based on the valid data content of adjacent grids. This process is continued until all missing parameters are filled, and all first parameter set content that has undergone anomaly removal and missing item filling is integrated. The integrated complete data is then uniformly composed of standardized soil sampling data. When sorting out the sampling data of scattered farmland in the southwestern mountainous region, based on the data content of the surrounding complete sampling grids, spatial interpolation is used to fill in the soil parameters corresponding to the scattered missing plots in the mountainous region, ultimately generating standardized soil sampling data for the area.

[0041] Step c1: Extract the second parameter set from the meteorological time series data. The second parameter set includes effective accumulated temperature, active accumulated temperature, effective precipitation, and active precipitation.

[0042] Furthermore, the meteorological time-series data consists of raw, time-series monitoring data collected continuously by meteorological monitoring stations at all levels over a long period. This data includes supplementary text unrelated to meteorological parameters, such as station equipment maintenance records and monitoring environment notes. The second parameter set is composed of four core meteorological statistical indicators: effective accumulated temperature, active accumulated temperature, effective precipitation, and active precipitation. These four indicators correspond to the statistical content of heat accumulation and water replenishment during the crop growth cycle, respectively. The entire text of the meteorological time-series data was traversed line by line, identifying the information attributes of different fields. Irrelevant content such as station maintenance records and instrument malfunction notes was removed. The time-series monitoring content corresponding to the four types of meteorological statistical indicators was extracted separately. All the selected indicator information was then uniformly summarized and collected, and the summarized complete information formed the second parameter set. During the compilation of the meteorological raw archives for the North China winter wheat planting area, the daily management notes of the stations were removed one by one, and the relevant time-series data of the four types of meteorological indicators were extracted separately to complete the construction of the second parameter set for the area.

[0043] Step c2: Assign administrative division codes to the second parameter set based on meteorological station data.

[0044] Furthermore, meteorological station data refers to the supporting original ledgers continuously collected from ground meteorological observation points deployed in different administrative regions. Administrative division codes are localized identifiers compiled based on unified regional division standards. The original data source for each second parameter set is bound to fixed meteorological station information. The textual information of the region to which each meteorological station belongs is retrieved one by one for each second parameter set. Based on the regional textual information, the corresponding standardized administrative division code is searched and matched. The retrieved code is then bound and appended to the end of the entry for this second parameter set. This process is repeated for all second parameter sets within the region, ensuring no valid meteorological parameter record is missed. When analyzing meteorological data from contiguous rice-growing areas in the Yangtze-Huaihe River Basin, the geographical scope of each meteorological station under the jurisdiction of each county is located one by one. Based on the geographical information, the corresponding administrative division code is matched for all second parameter sets generated by the stations.

[0045] Step c3: Classify and integrate the second parameter set according to the administrative division code to obtain standardized meteorological time series data.

[0046] Furthermore, all second parameter set entries with completed administrative division codes were retrieved. Using the attached administrative division codes as the classification benchmark, a collection process was conducted, grouping all second parameter set entries with the same administrative division code into the same storage group. After grouping, the generation time order of all time-series data within each group was analyzed, and the field layout format and document storage standards of each parameter within the group were uniformly adjusted. Formatting errors caused by scattered data entry were corrected. After all grouping and standardization work was completed, the standardized parameter content within all groups was integrated, and all integrated data constituted standardized meteorological time-series data. In the stage of summarizing meteorological data from the main summer maize producing areas of the Huang-Huai-Hai Plain, multi-source meteorological records were split according to the corresponding administrative division codes of districts and counties. After uniformly organizing and arranging the same code parameters, standardized meteorological time-series data for the region was generated.

[0047] Step d1: Obtain the pre-constructed standardized knowledge graph of diseases and pests. The standardized knowledge graph of diseases and pests includes disease name, pest name, pathogen, source of pests, stage of damage and control threshold.

[0048] Furthermore, the standardized knowledge graph for pests and diseases is a professional knowledge base pre-built according to plant protection industry standards. Disease names refer to the standardized names of various infectious diseases of crops, pest names refer to the standardized names of various phytophagous pests, pathogens refer to the definitions of the original microorganisms that induce crop diseases, pest sources refer to the source carriers of the initial reproduction and spread of pests, damage stages refer to different growth cycle intervals of crops affected by pests and diseases, and control thresholds refer to the baseline conditions for determining whether pest and disease control operations are necessary. The complete pre-built standardized knowledge graph for pests and diseases is retrieved, fully including the six categories of standardized definition information recorded within the graph, and retaining the graph's established pest and disease classification hierarchy and entry structure without altering the graph's built-in information classification logic during the retrieval process. When organizing data on pests and diseases of tropical fruits and vegetables in South China, the pre-built standardized knowledge graph for pests and diseases in the general version for grains, oils, fruits, and vegetables is directly retrieved, fully preserving all pest and disease baseline information within the graph.

[0049] Step d2 assigns administrative division codes to the unstructured data on pests and diseases.

[0050] Furthermore, the unstructured pest and disease data consists of field pest and disease survey texts collected from official monitoring documents and publicly available agricultural documents, lacking fixed standardized fields. Each original data record corresponds to a fixed field survey location. The textual content of the survey location accompanying each piece of unstructured pest and disease data was read one by one. Based on the location text matching standard administrative division codes, the matched administrative division codes were added as auxiliary fields to the end of the corresponding pest and disease text. The coding binding operation for all pest and disease texts was completed one by one. After all text coding configurations were completed, they were collected and archived for subsequent annotation processing. During the compilation of field pest and disease handbook data from the main citrus-producing areas in southern China, based on the village and town survey addresses recorded in the handbooks, the corresponding administrative division codes were matched for each individual survey text.

[0051] Step d3 involves labeling unstructured pest and disease data based on administrative division codes and a standardized knowledge graph of pests and diseases to obtain standardized unstructured pest and disease data.

[0052] Furthermore, the unstructured data on pests and diseases, each with its own administrative division code, was retrieved. The administrative division code of each entry pinpointed the geographical area corresponding to that pest or disease data entry. Based on the baseline naming and attribute information of various pests and diseases stored within the pre-constructed standardized knowledge graph of pests and diseases, the descriptions of pests and diseases in the original manuscript were compared word-for-word. The colloquial terms and unmarked descriptions within the manuscript were replaced with standardized terms built into the graph. Simultaneously, corresponding annotation fields were added based on the damage stages and control thresholds recorded in the graph. The annotation rules for all manuscripts were standardized. After all annotation work was completed, all annotated text data was compiled to form standardized unstructured data on pests and diseases. When compiling the research notes on pests and diseases in rice fields in the middle and lower reaches of the Yangtze River, the colloquial terms of pests and diseases within the manuscript were corrected by comparing with the standardized knowledge graph of pests and diseases. After supplementing the attribute annotations for each pest and disease, standardized pest and disease data for the region was generated.

[0053] Step e1: Extract the third parameter set from the seed basic approval data. The third parameter set includes variety name, approval number, approval area, suitable planting season, and characteristics.

[0054] Furthermore, the seed basic approval data originates from the original publicly released variety approval announcements. The announcement text includes irrelevant sections such as the public notice management regulations and application instructions. The third parameter set contains five key categories of variety registration information: variety name, approval number, approval region, suitable planting season, and characteristics. The variety name is the official name for the crop category; the approval number is the unique registration identifier issued by the government; the approval region is the legal planting area where the variety has passed trials and verification; the suitable planting season is the appropriate sowing period for the variety; and the characteristics are the inherent attributes related to the variety's growth and stress resistance. All original approval announcements were read through, and irrelevant text such as policy explanations and supplementary provisions were removed. The text corresponding to the five categories of registration information was extracted separately, and all the selected information was compiled to form the third parameter set. When reviewing the approval and public notice documents for summer maize in the Huang-Huai-Hai Plain over the years, the supplementary management clauses were removed, and the five core registration information categories were extracted to form the third parameter set for maize varieties in the region.

[0055] Step e2 assigns administrative division codes to the approved areas in the third parameter set.

[0056] Furthermore, the approved area refers to the administrative region within which the varieties recorded in the third parameter set are permitted to be planted. The geographical content is primarily represented by the textual description of the prefecture-level city or district / county name. The approved area textual content recorded in each entry within the third parameter set is retrieved one by one. Based on the administrative division coding standard resource, the exclusive coding content corresponding to the textual region is retrieved. The retrieved and matched administrative division codes are added to the corresponding subordinate field of the approved area for this entry. This coding binding operation is performed sequentially through all entries in the third parameter set, ensuring that the planting location of each variety registration information can be locked based on the code. During the stage of organizing the registration data for northern spring soybean varieties, based on the textual description of the planting prefecture in the entries, the corresponding administrative division codes are retrieved and matched one by one and attached to the approved area field.

[0057] Step e3: Identify and remove duplicate approval information from the third parameter set.

[0058] Furthermore, duplicate approval information refers to multiple highly overlapping registration entries for the same crop variety after multiple rounds of approval and public announcement. The variety name combined with the approval number can uniquely identify the registration status of a single crop variety. A horizontal comparison is performed on all entries within the entire third parameter set, simultaneously verifying the two key identifiers: the variety name and the approval number. If both identifiers are completely identical, the corresponding entry is determined to be duplicate approval information. The single registration document with the most complete content is retained, while other overlapping and redundant entries are directly eliminated. After the full comparison and cleanup of all entries, only the third parameter set without duplicate content remains. When screening domestic winter wheat variety registration documents, duplicate approval information resulting from multiple public announcements of the same variety is eliminated based on the variety name and approval number, retaining only the unique and valid registration entry.

[0059] Step e4: Classify and archive the third parameter set according to crop type to obtain standardized seed basic approval data.

[0060] Furthermore, crop types are standardized classification categories based on crop uses and category attributes, with grain crops, vegetable crops, and oilseed crops all belonging to different classification categories. All entries in the third parameter set, after duplicate information removal, are read. The crop category-related descriptions within each entry are extracted, and the entries are categorized under the corresponding crop type classification category based on these descriptions. All third parameter set data within the same category are collected and archived centrally. The storage arrangement and field formatting of all data within each category are standardized. After all classification and archiving operations are completed, the complete data within each category is integrated to generate standardized seed basic approval data. When collecting domestic fruit, vegetable, and grain variety registration data, they are archived according to their categories within the corresponding crop categories. After all data is organized, complete standardized seed basic approval data is generated.

[0061] Step S103: Obtain the first index and the second index, and associate the second multi-source heterogeneous agricultural data with the first index and the second index respectively to obtain the third multi-source heterogeneous agricultural data. The first index is the administrative division code index, and the second index is the crop type index.

[0062] Specifically, step S103 includes: Step f1: Obtain the first index and the second index based on the public platform.

[0063] Furthermore, the public platform refers to the official online public disclosure carriers for national standards and agricultural industry classification specifications, including the national standard standardization disclosure platform and agricultural industry data public release channels. The first index is the administrative division code index, which fully includes standardized code entries and coding specifications corresponding to administrative regions at all levels in China. The second index is the crop type index, which includes standardized classification entries and category division details corresponding to all categories of crops. Log in to various verified public platforms, locate the relevant specifications for administrative division codes and crop classification specifications in the platform's resource search interface, initiate resource download requests, and after downloading, check the format and classification content of each entry in the index, remove irrelevant attachments such as policy guides and clause explanations provided by the platform, organize and filter the pure index entries, and archive and store them uniformly to complete the complete acquisition of the first and second indexes. When organizing the supporting index resources for the Huang-Huai-Hai major grain-producing area, extract all entries of the administrative division code index from the national standard disclosure platform, download the complete content of the crop type index from the public platform of the agricultural authorities, remove redundant supplementary information, and archive and retain both types of index files.

[0064] Step f2 involves matching the standardized soil sampling data, standardized meteorological time-series data, standardized pest and disease unstructured data, and standardized seed basic approval data with the first index at the administrative division dimension.

[0065] Furthermore, standardized soil sampling data, standardized meteorological time-series data, standardized pest and disease unstructured data, and standardized seed basic approval data all have their corresponding administrative division codes attached during the pre-processing stage. The first index serves as the standard reference directory for administrative division codes across the entire region, and the directory uniformly defines the writing format and location correspondence of all codes. The administrative division code fields within each of the four types of standardized data are retrieved one by one. Based on the standard code entries stored within the first index, content comparison is performed item by item, correcting any discrepancies in the writing format of codes within each type of data. A new association marker is added to each data entry, recording the binding relationship between the data entry and the corresponding code entry within the first index. The comparison and marking operations for all data are completed in batches according to data categories, and all data sequentially completes the association matching process at the administrative division level. When sorting through the four types of standardized data for the rice planting areas in the middle and lower reaches of the Yangtze River, the administrative division code attached to each data entry is retrieved one by one, and the format is standardized by comparing it with the standard code content of the first index. Association markers are added synchronously, completing the association matching between the data of the entire region and the first index.

[0066] Step f3 involves matching the standardized soil sampling data, standardized meteorological time-series data, standardized pest and disease unstructured data, and standardized seed basic approval data (which have been correlated and matched at the administrative division level) with the second index at the crop type level to generate the third multi-source heterogeneous agricultural data.

[0067] Furthermore, the second index serves as a reference directory for defining classification standards for various crops. Within this directory, different crop classification entries are divided according to agricultural planting attributes. After the standardized data for each category completes the matching along the administrative division dimension, the entries retain relevant textual descriptions of the crops. Each of the four categories of standardized data already bound to the first index is retrieved. Descriptive text related to the crop category is extracted from the text of each individual data entry. Based on the pre-set standardized crop classification entries within the second index, the corresponding classification content is matched one by one. Crop classification association identifiers are added to the existing data entries. These identifiers record the binding relationship between the current data entry and the corresponding classification entry in the second index. This process of classification matching and identifier addition is repeated for all categories of multi-source data. After all entries complete the dual-index association, all data content is collected and integrated. The integrated complete data forms the third multi-source heterogeneous agricultural data. When processing data from the North China winter wheat planting area, the crop information recorded in each data entry is extracted, matched against the second index, and the classification identifiers are added. Then, the data from the entire area is summarized to generate the corresponding third multi-source heterogeneous agricultural data for that region.

[0068] Step S104: Obtain the first fusion strategy, and generate the first fusion dataset based on the first fusion strategy and the third multi-source heterogeneous agricultural data. The first fusion strategy is pre-built based on the first index and the second index.

[0069] Specifically, step S104 includes: Step g1: Obtain the first fusion strategy, which is a dual-keyword matching fusion rule.

[0070] Furthermore, the dual-keyword matching fusion rule refers to a data merging specification built upon two key identifiers: the administrative division code corresponding to the first index and the crop classification corresponding to the second index. The rule internally records the data collection order and field merging standards corresponding to different keywords. The complete first fusion strategy document within the pre-archived strategy resource directory is retrieved, and the grouping logic, matching order, and field merging details recorded within the document are checked line by line. After confirming the document content is complete and without omissions, all rule content is retained. Subsequent full-process data fusion operations execute the retained dual-keyword matching fusion rules throughout. For example, before conducting data fusion operations in the Huang-Huai-Hai grain production area, the dual-keyword matching fusion rule adapted to the grain and oil categories is retrieved, and the details related to regional grouping and crop matching within the rule are checked. After confirming the content is correct, the corresponding rule is set as the unified execution standard for regional data fusion.

[0071] Step g2: Based on the dual-keyword matching and fusion rules, the standardized soil sampling data, standardized meteorological time series data, and standardized pest and disease unstructured data in the third multi-source heterogeneous agricultural data are grouped and fused within groups using the first index and administrative division code to obtain several groups of first group data.

[0072] Furthermore, the first group of data refers to a collection of multiple types of environmental data gathered based on the same administrative division. Standardized soil sampling data, standardized meteorological time-series data, and standardized unstructured pest and disease data are uniformly classified as environmental data. All environmental data entries have been bound to the corresponding administrative division codes and the first index association marker. Based on the grouping logic recorded in the dual-keyword matching and fusion rules, the administrative division code content of all environmental data is extracted, and the three types of environmental data carrying the same administrative division code are gathered into the same group. After the grouping is completed, the field layout format of various types of data within the same group is uniformly adjusted. The scattered and independent single soil records, meteorological records, and pest and disease records are merged according to the rule requirements. All environmental data are continuously traversed to complete all grouping and intra-group integration, generating multiple sets of first group data in sequence. When collecting environmental data of the contiguous rice planting area in the Jianghuai region, soil, meteorological, and pest and disease data of the same district and county are collected based on the administrative division code. After the field structure is standardized, the first group of data exclusive to the corresponding district and county is generated.

[0073] Step g3: Based on the dual-keyword matching and fusion rule, the second index is used to match the crop type of each group of first group data to obtain the environmental adaptation data corresponding to each crop type.

[0074] Furthermore, the environmental adaptation data consists of aggregated environmental parameter information specific to a single crop category. The second index includes standardized classification entries for various crops, and each first group of data contains scattered environmental information covering all crop categories within the jurisdiction. Based on the crop selection details recorded in the dual-keyword matching and fusion rules, relevant descriptive information for various crops is extracted from the main text of each group of first group data. The extracted crop descriptions are then matched one by one with the standardized crop classification entries within the second index to lock all environmental records corresponding to each standard crop within a single group. Redundant environmental information irrelevant to the current crop classification within the group is removed, and the filtered and retained environmental data for the same crop is integrated and summarized. The summarized content is fixed as the environmental adaptation data for the corresponding crop type. When processing the grouped data of the contiguous spring maize planting area in Northeast China, the environmental information belonging to the spring maize category within the group is filtered based on the second index, other miscellaneous grain supporting data is removed, and the remaining content is integrated to form spring maize-specific environmental adaptation data.

[0075] Step g4 involves matching each crop type's environmental adaptation data and seed basic approval data item by item to obtain several integrated data entries, and then constructing the first fused dataset based on these integrated data entries.

[0076] Furthermore, the integrated data entry is a complete data record formed by merging the environmental parameters of a single variety with the variety registration parameters. After preprocessing, the seed basic approval data is fully labeled with the administrative division code and crop type corresponding to the approval area. Based on the entry splicing specifications recorded by the dual-keyword matching fusion rules, the administrative division information and crop classification information corresponding to each set of environmental adaptation data are retrieved one by one. At the same time, variety registration entries with the same administrative division code and the same crop classification are retrieved within the seed basic approval data. The various parameter information contained in the environmental adaptation data and the third parameter set contained in the seed basic approval data are merged and included in the same entry. The pairing and integration of all environmental adaptation data and corresponding variety information are completed iteratively, and all generated integrated data entries are summarized. After all entries are collected, a complete first fusion dataset is formed. When sorting out the data of citrus planting areas in the south, based on the citrus approval data and supporting environmental data of the same region and variety, integrated data entries are generated one by one and then summarized to generate the first fusion dataset of the area.

[0077] Step S105: Perform consistency check, integrity check and accuracy check on the first fused dataset in sequence to obtain the second fused dataset.

[0078] Specifically, step S105 includes: Step h1: Verify the logical matching relationship of the administrative division code, geographic coordinates and crop type fields in the first fused dataset, remove data entries with logical field conflicts and coordinate errors, and obtain the first fused dataset after consistency verification.

[0079] Furthermore, the administrative division code serves as a standardized code to indicate the geographical scope of the data; geographic coordinates are used to accurately locate the spatial position of the plot; and the crop type field records the standardized classification name of the corresponding crop. These three fields must maintain a logical correspondence between spatial location and the appropriate crop. Each entry in the first fused dataset is traversed sequentially. The contents of the administrative division code, geographic coordinates, and crop type fields are extracted from each entry. The inherent correspondence between these three fields is verified item by item to confirm that the spatial scope defined by the geographic coordinates and the administrative jurisdiction corresponding to the administrative division code are consistent. Simultaneously, it is confirmed that the labeled crop type belongs to a commonly introduced variety within the corresponding jurisdiction. During the screening process, all data entries with mismatched spatial scopes or incorrect coordinate entries are separately identified and removed. All remaining compliant entries are summarized and collected, forming the first fused dataset after consistency verification. For example, when conducting the verification of the Huang-Huai-Hai winter wheat fused data, each data entry's geographical code, plot coordinates, and registered crop type are verified, and entries with logical mismatches are removed.

[0080] Step h2 involves using the mean data from similar regions and model inference to complete the missing core parameter fields within the data entries of the first fused dataset after consistency verification. The core parameter fields include the first parameter set, the second parameter set, the pest and disease risk parameters, and the third parameter set, resulting in the first fused dataset after integrity verification.

[0081] Furthermore, the core parameter fields contain four categories of fixed parameters: the first parameter set corresponds to soil physicochemical indicators; the second parameter set corresponds to meteorological statistics; pest and disease risk parameters are generated based on pest and disease annotations; and the third parameter set records variety approval and registration information. Each data entry within the first fused dataset after consistency verification is reviewed, and the completion status of each core parameter field within each entry is checked, marking the categories of parameters with missing content. When the mean data from similar regions is selected for completion, multiple data entries with complete parameters under the same administrative division and crop classification are retrieved, and the complete parameter information is extracted to fill in the missing fields. When the model inference completion method is selected, inference logic is built based on the existing complete parameters in the same region, and the missing parameter content is filled based on the inference results. All missing fields are processed iteratively according to the two completion paths. After all missing items are filled, all entry content is integrated to form the first fused dataset after integrity verification.

[0082] Step h3: Perform accuracy verification on the first fused dataset after integrity verification to generate the second fused dataset. Accuracy verification includes sampling comparison verification from authoritative agricultural databases and manual sampling verification by agricultural experts.

[0083] Furthermore, the authoritative agricultural database sampling comparison and verification refers to the content benchmarking operation based on standard measured data retained in the industry's official archived database, while the manual sampling verification by agricultural experts refers to hiring professionals in the fields of breeding and plant protection to manually verify the authenticity of the sampled items. Multiple sampling batches are divided, and a specified number of data items are randomly selected from the first fused dataset after integrity verification. For the selected items, the authoritative agricultural database sampling comparison and verification is initiated first. Based on the standard archived information of the same region and crop within the authoritative database, each item's recorded parameters are compared, and items whose content deviates from the standard records are marked. After completing the database comparison, manual sampling verification by agricultural experts continues. Professional experts review the variety information, pest and disease records, and environmental parameters within each sampled item, correcting any input errors discovered during the manual verification process. After all batches of sampling verification are completed, all verified and corrected data are collected, and the collected complete data is uniformly generated into a second fused dataset.

[0084] Step S106: Standardize the assets of the second fused dataset to generate a standardized data access interface and an agricultural variety adaptation dataset.

[0085] Specifically, standardized asset encapsulation refers to the process of standardizing the field formats, storage directory structures, and resource organization of the second fused dataset after consistency, integrity, and accuracy verification. Standardized data access interfaces are interactive ports that provide compliant data retrieval channels for various agricultural technology business procedures. Agricultural variety adaptation datasets are encapsulated, finished data resources that can be directly used for introduction trials and regional variety layout analysis. The process involves systematically reviewing all integrated data entries within the second fused dataset, standardizing the field naming rules and text formatting for each entry's first parameter set, second parameter set, pest and disease risk parameters, and third parameter set. A multi-level storage directory is built based on administrative division codes and crop types, and all data is stored sequentially in the corresponding storage partitions according to the directory classification. Based on the standardized directory structure, a standardized data access interface is developed, clarifying the parameter range and data output format that the interface can retrieve. After interface configuration, all standardized data is summarized, and the summarized content constitutes the final agricultural variety adaptation dataset. When processing the second fused dataset of summer maize production area in Huang-Huai-Hai Plain, the format of all data fields in the region was unified, and the storage catalog was based on the district and county divisions and grain varieties. A dedicated query interface was developed to produce a summer maize agricultural variety adaptation dataset for use by local agricultural technology extension departments.

[0086] The method for constructing an agricultural variety adaptation dataset based on multi-source heterogeneous data provided in this embodiment firstly covers the core data dimensions involved in the crop growth adaptation process, such as soil and water environment, climate conditions, pest and disease impacts, and variety approval qualifications, by collecting soil sampling data, meteorological time series data, unstructured pest and disease data, and seed basic approval data. This fully covers the basic data system required for agricultural variety adaptation assessment. The multi-dimensional data acquisition model fully meets the underlying data needs of agricultural variety adaptation scenarios, avoiding the information gaps caused by single data dimensions. Based on multi-dimensional raw data, it provides complete and business-scenario-appropriate raw data support for subsequent data standardization processing, data association and fusion, and other full-process technical operations. Secondly, it adopts differentiated processing strategies to carry out targeted standardization processing of multi-source heterogeneous agricultural data. Based on the structural characteristics, data forms, and information carrier differences of different types of data, it configures corresponding data regularization logic. For structured soil and water data, meteorological data, and seed data, it performs parameter extraction and regularization processing, and for unstructured pest and disease data, it performs annotation and standardization sorting. It adapts to the own attribute characteristics of various heterogeneous data, solves the technical problems of inconsistent formats, disordered parameters, and mismatched information regularization logic of different data sources, realizes the standardized conversion of various raw heterogeneous data, and unifies the basic form of multi-source data. Then, by configuring two standardized index systems—administrative division code index and crop type index—two-way dimensional association and binding operations are carried out on the standardized multi-source agricultural data. A unified data association benchmark is built based on these two dedicated indexes, assigning unified spatial attribution and crop classification identifiers to all types of agricultural data, and establishing an underlying association logic that enables interoperability and matching between heterogeneous data across categories. The standardized index configuration mode breaks down the association barriers formed by the independent storage and organization of various types of data, providing a technical foundation for linked matching of scattered and independent multi-source data, achieving orderly aggregation and dimensional unification of multi-source heterogeneous data. Subsequently, by pre-constructing a dedicated first fusion strategy based on the dual index system, the fusion process of multi-source data is coordinated based on standardized dual-keyword matching logic, abandoning the disorderly splicing processing mode. Data grouping, adaptive data filtering, and variety information matching operations are completed based on the index dimension. Environmental data aggregation and variety data pairing are strictly completed according to the dual logic of spatial and crop dimensions, standardizing the overall execution logic of multi-source data fusion. An orderly data fusion model enables the orderly transformation of fragmented single-category data into integrated data, ensuring the logic, standardization, and consistency of the data fusion process, and avoiding logical confusion and information mismatch problems in the data fusion process.Furthermore, through progressive quality control operations involving consistency, integrity, and accuracy checks on the initial fusion dataset, data quality is optimized layer by layer from multiple technical levels, including field logical relationships, completeness of core parameters, and overall data authenticity. This process systematically addresses and corrects internal logical conflicts, fills in missing core parameters, and standardizes overall data accuracy. A systematic data quality control logic is built based on multi-level verification methods, and this progressive verification model comprehensively optimizes the internal data structure and content quality of the dataset. Finally, standardized asset encapsulation is performed on the fusion dataset after quality verification. This unifies the dataset's field definitions, storage formats, and data specifications, establishes a multi-dimensional retrieval system, and generates standardized data access interfaces. The standardized asset encapsulation model enables the compliant transformation of fragmented data resources into standardized, structured digital assets, establishes a unified data retrieval and calling mechanism, standardizes the output format and usage of the dataset, and achieves standardized storage, convenient retrieval, and stable reuse of agricultural variety adaptation datasets based on a standardized interface system. By implementing this solution, we have solved the problem that agricultural variety-related data processing technologies suffer from multiple technical shortcomings, such as fragmented data standardization rules, lack of unified index support for cross-source association, crude fusion logic, and lack of a data quality control system, making it difficult to produce standardized and usable agricultural variety-adapted datasets.

[0087] The above are embodiments of the method for constructing an agricultural variety adaptation dataset based on multi-source heterogeneous data provided in this application. Other embodiments of constructing an agricultural variety adaptation dataset based on multi-source heterogeneous data provided in this application are described below.

[0088] This invention also discloses a system for constructing agricultural variety adaptation datasets based on multi-source heterogeneous data, such as... Figure 2 As shown, the system includes: The acquisition module is used to acquire the first multi-source heterogeneous agricultural data, which includes soil sampling data, meteorological time-series data, unstructured pest and disease data, and seed basic approval data. The processing module is used to standardize the first multi-source heterogeneous agricultural data using a differentiated processing strategy to obtain the second multi-source heterogeneous agricultural data. The second multi-source heterogeneous agricultural data includes standardized soil sampling data, standardized meteorological time-series data, standardized unstructured pest and disease data, and standardized seed basic approval data. The association module is used to obtain the first index and the second index, and associate the second multi-source heterogeneous agricultural data with the first index and the second index respectively to obtain the third multi-source heterogeneous agricultural data. The first index is the administrative division code index, and the second index is the crop type index. The fusion module is used to obtain a first fusion strategy and generate a first fusion dataset based on the first fusion strategy and third multi-source heterogeneous agricultural data. The first fusion strategy is pre-built based on a first index and a second index. The verification module is used to sequentially perform consistency verification, integrity verification, and accuracy verification on the first fused dataset to obtain the second fused dataset. The generation module is used to standardize the asset encapsulation of the second fused dataset, and generate standardized data access interfaces and agricultural variety adaptation datasets.

[0089] The agricultural variety adaptation dataset construction system provided in this embodiment firstly covers the core data dimensions involved in the crop growth adaptation process, such as soil and water environment, climate conditions, pest and disease impact, and variety approval qualifications, by collecting soil sampling data, meteorological time series data, unstructured pest and disease data, and seed basic approval data. It fully covers the basic data system required for agricultural variety adaptation assessment. The multi-dimensional data acquisition model fully meets the underlying data needs of agricultural variety adaptation scenarios, avoiding the information gaps caused by single data dimensions. Based on multi-dimensional raw data, it provides complete and business-scenario-appropriate raw data support for subsequent data standardization processing, data association and fusion, and other full-process technical operations. Secondly, it adopts differentiated processing strategies to carry out targeted standardization processing of multi-source heterogeneous agricultural data. Based on the structural characteristics, data forms, and information carrier differences of different types of data, it configures corresponding data regularization logic. For structured soil and water data, meteorological data, and seed data, it performs parameter extraction and regularization processing, and for unstructured pest and disease data, it performs annotation and standardization sorting. It adapts to the own attribute characteristics of various heterogeneous data, solves the technical problems of inconsistent formats, disordered parameters, and mismatched information regularization logic of different data sources, realizes the standardized conversion of various raw heterogeneous data, and unifies the basic form of multi-source data. Then, by configuring two standardized index systems—administrative division code index and crop type index—two-way dimensional association and binding operations are carried out on the standardized multi-source agricultural data. A unified data association benchmark is built based on these two dedicated indexes, assigning unified spatial attribution and crop classification identifiers to all types of agricultural data, and establishing an underlying association logic that enables interoperability and matching between heterogeneous data across categories. The standardized index configuration mode breaks down the association barriers formed by the independent storage and organization of various types of data, providing a technical foundation for linked matching of scattered and independent multi-source data, achieving orderly aggregation and dimensional unification of multi-source heterogeneous data. Subsequently, by pre-constructing a dedicated first fusion strategy based on the dual index system, the fusion process of multi-source data is coordinated based on standardized dual-keyword matching logic, abandoning the disorderly splicing processing mode. Data grouping, adaptive data filtering, and variety information matching operations are completed based on the index dimension. Environmental data aggregation and variety data pairing are strictly completed according to the dual logic of spatial and crop dimensions, standardizing the overall execution logic of multi-source data fusion. An orderly data fusion model enables the orderly transformation of fragmented single-category data into integrated data, ensuring the logic, standardization, and consistency of the data fusion process, and avoiding logical confusion and information mismatch problems in the data fusion process.Furthermore, through progressive quality control operations involving consistency, integrity, and accuracy checks on the initial fusion dataset, data quality is optimized layer by layer from multiple technical levels, including field logical relationships, completeness of core parameters, and overall data authenticity. This process systematically addresses and corrects internal logical conflicts, fills in missing core parameters, and standardizes overall data accuracy. A systematic data quality control logic is built based on multi-level verification methods, and this progressive verification model comprehensively optimizes the internal data structure and content quality of the dataset. Finally, standardized asset encapsulation is performed on the fusion dataset after quality verification. This unifies the dataset's field definitions, storage formats, and data specifications, establishes a multi-dimensional retrieval system, and generates standardized data access interfaces. The standardized asset encapsulation model enables the compliant transformation of fragmented data resources into standardized, structured digital assets, establishes a unified data retrieval and calling mechanism, standardizes the output format and usage of the dataset, and achieves standardized storage, convenient retrieval, and stable reuse of agricultural variety adaptation datasets based on a standardized interface system. By implementing this solution, we have solved the problem that agricultural variety-related data processing technologies suffer from multiple technical shortcomings, such as fragmented data standardization rules, lack of unified index support for cross-source association, crude fusion logic, and lack of a data quality control system, making it difficult to produce standardized and usable agricultural variety-adapted datasets.

[0090] The agricultural variety adaptation dataset construction system based on multi-source heterogeneous data in this embodiment is presented in the form of functional units. Here, a unit refers to an ASIC (Application Specific Integrated Circuit), a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

Claims

1. A method for constructing an agricultural variety adaptation dataset based on multi-source heterogeneous data, characterized in that, The method includes: Acquire the first multi-source heterogeneous agricultural data, which includes soil sampling data, meteorological time-series data, unstructured pest and disease data, and seed basic approval data; The first multi-source heterogeneous agricultural data is standardized using a differentiated processing strategy to obtain the second multi-source heterogeneous agricultural data. The second multi-source heterogeneous agricultural data includes standardized soil sampling data, standardized meteorological time-series data, standardized unstructured pest and disease data, and standardized seed basic approval data. Obtain the first index and the second index, and associate the second multi-source heterogeneous agricultural data with the first index and the second index respectively to obtain the third multi-source heterogeneous agricultural data. The first index is the administrative division code index, and the second index is the crop type index. A first fusion strategy is obtained, and a first fusion dataset is generated based on the first fusion strategy and the third multi-source heterogeneous agricultural data. The first fusion strategy is pre-built based on the first index and the second index. The first fused dataset is subjected to consistency verification, integrity verification, and accuracy verification in sequence to obtain the second fused dataset. The second fused dataset is encapsulated into standardized assets to generate standardized data access interfaces and agricultural variety adaptation datasets.

2. The method according to claim 1, characterized in that, The acquisition of the first multi-source heterogeneous agricultural data includes: Soil sampling data was obtained based on the national soil survey database; Meteorological time-series data are obtained based on the public interface of the National Meteorological Center and third-party meteorological service platforms; Unstructured pest and disease data were obtained based on pest and disease forecasting data from the Ministry of Agriculture and Rural Affairs and publicly available online information. Basic seed approval data was obtained from the variety approval announcements of the National Agricultural Technology Extension Service Center.

3. The method according to claim 2, characterized in that, The standardization process of the first multi-source heterogeneous agricultural data using a differentiated processing strategy to obtain the second multi-source heterogeneous agricultural data includes: Extract a first set of parameters from the soil sampling data, the first set of parameters including pH value, organic matter content, total nitrogen content, available phosphorus content and available potassium content; Map the discrete sampling point data corresponding to the first parameter set to the corresponding geographic grid; Assign administrative division codes to each geographic grid; The IQR outlier detection algorithm is used to remove outlier sampled values ​​from the first parameter set; Standardized soil sampling data are obtained by filling in the missing values ​​of the first parameter set in each geographic grid using spatial interpolation.

4. The method according to claim 3, characterized in that, The standardization process of the first multi-source heterogeneous agricultural data using a differentiated processing strategy to obtain the second multi-source heterogeneous agricultural data includes: Extract the second parameter set from the meteorological time series data. The second parameter set includes effective accumulated temperature, active accumulated temperature, effective precipitation, and active precipitation. Administrative division codes are assigned to the second parameter set based on meteorological station data; The second parameter set is classified and integrated according to administrative division codes to obtain standardized meteorological time series data.

5. The method according to claim 4, characterized in that, The standardization process of the first multi-source heterogeneous agricultural data using a differentiated processing strategy to obtain the second multi-source heterogeneous agricultural data includes: Obtain a pre-constructed standardized knowledge graph of diseases and pests, which includes disease name, pest name, pathogen, source of pests, stage of damage, and control threshold. Assign administrative division codes to the aforementioned unstructured pest and disease data; The unstructured data of pests and diseases is labeled based on administrative division codes and standardized knowledge graphs of pests and diseases to obtain standardized unstructured data of pests and diseases.

6. The method according to claim 5, characterized in that, The standardization process of the first multi-source heterogeneous agricultural data using a differentiated processing strategy to obtain the second multi-source heterogeneous agricultural data includes: Extract the third parameter set from the seed basic approval data. The third parameter set includes variety name, approval number, approval area, suitable planting season and characteristics. Assign administrative division codes to the approved areas in the third parameter set; Identify and remove duplicate approval information from the third parameter set; The third parameter set is classified and archived according to crop type to obtain standardized seed basic approval data.

7. The method according to claim 6, characterized in that, The step of obtaining the first index and the second index, and associating the second multi-source heterogeneous agricultural data with the first index and the second index respectively to obtain the third multi-source heterogeneous agricultural data includes: Obtain the first and second indexes based on a public platform; Standardized soil sampling data, standardized meteorological time-series data, standardized pest and disease unstructured data, and standardized seed basic approval data are respectively matched with the first index at the administrative division dimension. The standardized soil sampling data, standardized meteorological time-series data, standardized pest and disease unstructured data, and standardized seed basic approval data, which have been correlated and matched at the administrative division level, are respectively correlated and matched at the crop type level with the second index to generate the third multi-source heterogeneous agricultural data.

8. The method according to claim 7, characterized in that, The step of obtaining the first fusion strategy and generating a first fused dataset based on the first fusion strategy and the third multi-source heterogeneous agricultural data includes: Obtain the first fusion strategy, which is a dual-keyword matching fusion rule; Based on the dual-keyword matching and fusion rules, the standardized soil sampling data, standardized meteorological time series data, and standardized pest and disease unstructured data in the third multi-source heterogeneous agricultural data are grouped and fused within groups using the first index and administrative division code to obtain several groups of first group data; Based on the dual-keyword matching and fusion rule, the second index is used to match crop types in the first group of data in each group to obtain the environmental adaptation data corresponding to each crop type. Based on the environmental adaptation data and seed basic approval data corresponding to each crop type, several integrated data entries are obtained by matching them item by item. The first fused dataset is constructed based on these integrated data entries.

9. The method according to claim 8, characterized in that, The process of sequentially performing consistency checks, integrity checks, and accuracy checks on the first fused dataset to obtain the second fused dataset includes: Verify the logical matching relationship of the administrative division code, geographic coordinates and crop type fields in the first fused dataset, remove data entries with logical field conflicts and coordinate errors, and obtain the first fused dataset after consistency verification; The missing core parameter fields inside the data entries in the first fused dataset after consistency verification are completed by using the mean data of similar regions and the model inference completion method. The core parameter fields include the first parameter set, the second parameter set, the pest and disease risk parameters and the third parameter set, to obtain the first fused dataset after integrity verification. An accuracy check is performed on the first fused dataset after integrity verification to generate a second fused dataset. The accuracy check includes sampling comparison verification from authoritative agricultural databases and manual sampling verification by agricultural experts.

10. A system for constructing an agricultural variety adaptation dataset based on multi-source heterogeneous data, characterized in that, The system includes: The acquisition module is used to acquire the first multi-source heterogeneous agricultural data, which includes soil sampling data, meteorological time-series data, unstructured pest and disease data, and seed basic approval data. The processing module is used to standardize the first multi-source heterogeneous agricultural data using a differentiated processing strategy to obtain the second multi-source heterogeneous agricultural data. The second multi-source heterogeneous agricultural data includes standardized soil sampling data, standardized meteorological time-series data, standardized unstructured pest and disease data, and standardized seed basic approval data. The association module is used to obtain the first index and the second index, and associate the second multi-source heterogeneous agricultural data with the first index and the second index respectively to obtain the third multi-source heterogeneous agricultural data. The first index is an administrative division code index, and the second index is a crop type index. The fusion module is used to obtain a first fusion strategy and generate a first fusion dataset based on the first fusion strategy and the third multi-source heterogeneous agricultural data. The first fusion strategy is pre-built based on the first index and the second index. The verification module is used to sequentially perform consistency verification, integrity verification and accuracy verification on the first fused dataset to obtain the second fused dataset; The generation module is used to standardize the asset encapsulation of the second fused dataset, and generate standardized data access interfaces and agricultural variety adaptation datasets.