A Tropical Agricultural Resource Database Construction System Based on Agricultural Big Data

By constructing a tropical agricultural resource database and employing data preprocessing, multidimensional labeling, and genealogical management, the problems of loose organization of tropical agricultural data and lack of constraints in correlation analysis were solved, achieving efficient and accurate data management and analysis, and ensuring continuous improvement in data quality.

CN122086866APending Publication Date: 2026-05-26HAINAN QINGFENG BIOLOGICAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HAINAN QINGFENG BIOLOGICAL TECH CO LTD
Filing Date
2026-02-10
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing agricultural big data processing systems cannot accurately reflect the complex logical relationships between crop types, regional climates, and production processes when processing tropical agricultural data. This results in loose data organization, making it difficult to form a structured knowledge system. Furthermore, cross-data source correlation analysis lacks business scenario constraints, leading to low reliability and poor practicality of the analysis results.

Method used

The system for constructing a tropical agricultural resource database based on agricultural big data establishes a tropical agricultural resource genealogy framework through data preprocessing, multidimensional labeling, genealogy management, and correlation analysis. It also performs data verification and supplementation to ensure that data associations are conducted within semantically consistent and business-related data groups. The system uses preset rules to identify data gaps and trigger targeted supplementary collection.

Benefits of technology

It improves the efficiency of data retrieval, location, and overall management; enhances the internal logic and domain relevance of data organization; ensures the accuracy and reliability of correlation analysis results; and achieves efficient and accurate data correlation and continuous data quality maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122086866A_ABST
    Figure CN122086866A_ABST
Patent Text Reader

Abstract

This invention relates to the field of agricultural information technology, specifically to a system for constructing a tropical agricultural resource database based on agricultural big data. The system includes: a data preprocessing module, a data tagging module, a genealogy management module, a correlation analysis module, and a data verification and supplementation module. By tagging data in multiple dimensions according to tropical crop types, geographical and climatic zones, and production management stages, and mapping the tagged data to corresponding hierarchical nodes in a pre-defined tropical agricultural resource genealogy framework, the system achieves structured reorganization of data according to domain knowledge. Based on this, the system performs cross-data source correlation analysis within the same node of the genealogy framework, establishing precise data correlation links. This solution solves the problems of loose organization of tropical agricultural data and low accuracy of cross-source correlation, effectively improving the structured integration of data resources and the credibility of data analysis results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of agricultural information technology, and in particular to a system for constructing a tropical agricultural resource database based on agricultural big data. Background Technology

[0002] Currently, in the tropical agriculture sector, the widespread adoption of technologies such as the Internet of Things (IoT) and remote sensing has generated massive amounts of diverse and heterogeneous raw data, covering various aspects including crop information, environmental monitoring, and production operations. Existing agricultural big data processing systems generally employ common data cleaning, labeling, classification, and association analysis techniques to integrate these resources. A common approach is to standardize the data, assign broad labels, and then perform association rule mining and data fusion across the entire database based on a global data warehouse or using general algorithms.

[0003] These conventional technical solutions have shortcomings when processing tropical agricultural data, which has strong domain-specific characteristics. General labeling systems cannot accurately reflect the complex professional logical relationships between crop types, regional climates, and production stages, resulting in loose data organization and difficulty in forming a structured knowledge system consistent with agricultural science. Simultaneously, cross-data source global correlation analysis lacks necessary business scenario constraints, easily generating numerous invalid or even erroneous correlations between unrelated data, leading to low reliability and poor practicality of the analysis results, and failing to support refined agricultural decision-making. Therefore, how to deeply integrate tropical agricultural expertise into the data organization process and achieve accurate and efficient data correlation within this framework is a core issue that urgently needs to be addressed in building a high-quality agricultural resource database. Summary of the Invention

[0004] The purpose of this invention is to address the shortcomings of existing technologies by proposing a tropical agricultural resource database construction system based on agricultural big data.

[0005] To achieve the above objectives, the present invention adopts the following technical solution: a tropical agricultural resource database construction system based on agricultural big data, comprising: The data preprocessing module collects various raw data records from the field of tropical agriculture, performs data format standardization processing on the raw data records, and forms a standardized data set. The data tagging module performs multi-dimensional tagging on standardized datasets based on tropical crop types, geographical and climatic zones, and production management processes, generating initial agricultural resource data with classification tags. The genealogy management module constructs a tropical agricultural resource genealogy framework, mapping initial agricultural resource data with classification labels to the corresponding hierarchical nodes of the tropical agricultural resource genealogy framework according to their classification labels. The correlation analysis module performs data correlation analysis on initial agricultural resource data at the same level node within the tropical agricultural resource genealogy framework, and establishes data correlation links across data sources. The data verification and supplementation module verifies agricultural resource data organized through data association links according to preset data integrity rules. Data that fails verification is marked as data verification gaps, and the data supplementation and collection process is triggered.

[0006] Preferably, the multi-dimensional labeling of standardized datasets based on tropical crop species, geographical and climatic zones, and production management processes includes the following steps: Parse each data record in the standardized dataset and extract the entity object name and attribute description information contained in the data record; Match the entity object name with the tropical crop species name database and assign crop species labels to the data records; Analyze the spatial location keywords and time period keywords in the attribute description information, compare them with the geographic climate zoning database and the historical production season database respectively, and assign geographic area labels and time period labels to the data records. Identify specific agricultural operations or management measures terms involved in the attribute description information, compare them with the preset production link list, and assign production link labels to the data records; By combining crop type labels, geographical region labels, time period labels, and production process labels, corresponding classification labels for data records are generated.

[0007] Preferably, constructing a phylogenetic framework for tropical agricultural resources includes the following steps: The root node of the phylogenetic framework is defined as tropical agricultural resource data; Under the root node, establish first-level branch nodes according to the major categories of tropical crops; Under each primary branch node, secondary branch nodes are established based on different crop varieties or strains; Under the secondary branch nodes, tertiary branch nodes are established according to the different phenological stages within the crop growth cycle; Under the third-level branch nodes, fourth-level branch nodes are established according to different geographical and climatic zones; Under the fourth-level branch node, fifth-level branch nodes are established according to the types of resource data. The types of resource data include environmental monitoring data, soil property data, pest and disease occurrence data, agronomic operation record data, and yield and quality data.

[0008] Preferably, mapping initial agricultural resource data with classification labels to corresponding hierarchical nodes in the tropical agricultural resource genealogy framework according to their classification labels includes the following steps: Read the classification labels carried by the initial agricultural resource data; Match the crop type labels in the classification tags with the names of the first-level and second-level branch nodes in the tropical agricultural resource genealogy framework to determine the target second-level branch node; Match the time period label in the category label with the name of the third-level branch node under the target second-level branch node to determine the target third-level branch node; Match the geographic region label in the category label with the name of the fourth-level branch node under the target third-level branch node to determine the target fourth-level branch node; The subject content of the initial agricultural resource data is matched with the data type description of the fifth-level branch node under the target fourth-level branch node to determine the final target fifth-level branch node for storage, and the initial agricultural resource data is stored in it.

[0009] Preferably, establishing a data link across data sources includes the following steps: Under the same target level 5 branch node, initial agricultural resource data from different data collection sources are filtered out; Analyze the coupling relationship between initial agricultural resource data in terms of time period labels, geographic region labels, and specific data content; For initial agricultural resource data that have temporal continuity, spatial overlap, or complementary content, create association indexes between them; Based on the type of associated index, construct data traceability links, data spatial overlay links, or data content complementary links, and store the link information in the metadata layer of the tropical agricultural resource genealogy framework.

[0010] Preferably, verifying agricultural resource data organized via data association links according to preset data integrity rules includes the following steps: A data integrity template is defined for each fifth-level branch node in the tropical agricultural resource genealogy framework. The data integrity template specifies the data fields, data time span, and data spatial coverage density that should be included under the ideal state of the fifth-level branch node. Retrieve all the initial agricultural resource data actually stored under a certain fifth-level branch node and the data association links established thereunder; The actual data's field coverage, time span, and spatial density are compared item by item with the data integrity template. When the actual data fails to meet the threshold set by the data integrity template in any aspect such as field coverage, time span, or spatial density, it is determined that there is a data verification gap in the fifth-level branch node, and the specific type of the gap is recorded.

[0011] Preferably, the process of triggering supplementary data collection includes the following steps: Based on the specific type of data verification gap, a structured data supplementation task description is generated. The data supplementation task description clearly specifies the data subject to be supplemented, the target time range, the target geographical area, and the required data fields. Publish the data supplementation task description to the preset data acquisition task queue; The data acquisition task queue automatically matches devices or data source interfaces with corresponding data acquisition capabilities based on the task description; Send data acquisition instructions to the successfully matched device or interface, and receive the supplementary acquisition data returned. The supplementary collected data is formatted and multidimensionally labeled to form new initial agricultural resource data, which is then mapped to fill the original five-level branch nodes where data verification gaps existed.

[0012] Preferably, after mapping and filling the data inconsistencies in the original fifth-level branch nodes, the method further includes an associated link update step: Examine the classification labels and specific details of the new initial agricultural resource data; Under the same fifth-level branch node, the correlation between new initial agricultural resource data and existing initial agricultural resource data is automatically calculated; If new relationships are discovered, a new data link is established between the new initial agricultural resource data and the relevant existing initial agricultural resource data, or the existing data link is extended and expanded. Update the link information stored in the metadata layer of the tropical agricultural resource genealogy framework.

[0013] Preferably, before standardizing the data format of the original data records, a data credibility pre-screening step is also included: Each original data record is assigned a source credibility score, which is calculated based on the data provider's historical data quality rating, the calibration status of the data acquisition equipment, and the integrity of the data transmission link. Set a credibility threshold to filter out original data records whose source credibility score is lower than the credibility threshold; Subsequent format standardization processing is performed only on the original data records that have passed the credibility screening.

[0014] Preferably, the spatial location keywords and time period keywords in the analytical attribute description information are compared with the geographic climate zoning database and the historical production season database, respectively, and the data records are assigned geographic area labels and time period labels, including the following steps: Natural language parsing is performed on the attribute description information of each data record in the standardized dataset to identify and extract proper noun phrases describing geographical location and adverbial phrases describing time period as keywords; The extracted spatial location keywords are fuzzily matched with the administrative division names, topographic terms and climate zone classification names in the geographic climate zoning database, and the semantic similarity between each keyword and the database entry is calculated. When the semantic similarity exceeds the preset matching threshold, the standard geographic region code corresponding to the entry in the geographic climate zoning database is assigned to the data record to form a geographic region label; The extracted time period keywords are aligned with the crop phenological calendar, traditional agricultural solar terms, and planting season division information recorded in the historical production season database. Identify the specific production stage in the historical production season database to which the time period indicated by the time cycle keyword belongs, and assign a unique identifier to the data record for the specific production stage to form a time stage label.

[0015] Compared with the prior art, the advantages and positive effects of the present invention are as follows: Standardized data is tagged based on three dimensions: tropical crop type, geographical and climatic zone, and production management stage. These tags then automatically map the data to corresponding hierarchical nodes within a pre-constructed tropical agricultural resource hierarchy framework. Discrete data points are deeply and systematically reorganized according to the professional classification logic of the agricultural field. Data is no longer simply stacked according to keywords, but rather placed within a hierarchical topology defined by domain knowledge. This improves the efficiency of data retrieval, location, and overall management, and enhances the inherent logic and domain relevance of the data organization.

[0016] Based on the completed phylogenetic mapping, association analysis is strictly limited to nodes within the same phylogenetic hierarchy. That is, association analysis is performed only on subsets of data with the same crop, climate, and production stage labels, and cross-source data links are established. This technical solution effectively avoids noise interference from global association. It ensures that association analysis is always conducted within semantically consistent and business-relevant data groups, improving the accuracy and business relevance of the discovered data relationships, making the data analysis results more reliable and directly usable.

[0017] The system performs integrity checks on the data network formed through the aforementioned hierarchical organization and intra-node associations according to preset rules, and triggers targeted supplementary data collection for identified gaps. This mechanism ensures that data quality maintenance and data growth are no longer blind and passive. Based on existing and interconnected knowledge structures, the system can intelligently identify weak links or missing parts of the current data system, thereby driving purposeful and efficient new data collection and ensuring that the database maintains a high-value and structured complete state throughout its continuous evolution. Attached Figure Description

[0018] Figure 1 This is a timeline diagram of the tropical agricultural resource database construction system based on agricultural big data as described in this invention. Figure 2 Flowchart for constructing a phylogenetic framework for tropical agricultural resources; Figure 3 A flowchart for establishing data association links across data sources; Figure 4 A bar chart comparing the completeness of coffee environmental monitoring data before and after supplementary data collection; Figure 5 This is a graph used to evaluate the effectiveness of data verification and supplementation. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0020] In the description of this invention, it should be understood that the terms "length," "width," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships, are based on the orientation or positional relationships shown in the accompanying drawings and are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, in the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0021] See Figure 1The system first uses a data preprocessing module to collect various raw data records from the tropical agricultural sector and standardizes these records to form a standardized dataset. Then, a data tagging module tags this standardized dataset based on multiple dimensions, including tropical crop type, geographical climate zone, and production management stage, generating corresponding classification tags for each data record, thus forming initial agricultural resource data with classification tags. A genealogy management module pre-constructs a hierarchical tropical agricultural resource genealogy framework and maps these initial agricultural resource data with classification tags to corresponding hierarchical nodes within this framework. A correlation analysis module performs data correlation analysis on initial agricultural resource data at the same hierarchical node within the tropical agricultural resource genealogy framework, thereby establishing data correlation links across data sources. Finally, a data verification and supplementation module verifies the agricultural resource data organized through the data correlation links according to preset data integrity rules. Data that fails verification is marked as a data verification gap, triggering a data supplementation collection process to fill the gap.

[0022] In one embodiment of this invention, a data tagging module performs multi-dimensional tagging processing on a standardized dataset. The module parses each data record in the standardized dataset, extracting the entity name and attribute description information contained within the data record. Taking a data record from a tropical agriculture monitoring station as an example, the record contains the entity name "rubber tree" and the attribute description "rubber forests located in Danzhou City, Hainan Province, underwent rubber tapping during the 2023 rainy season." The data tagging module matches the entity name "rubber tree" with a pre-established tropical crop species name database, which contains the term "rubber tree." Entries such as "coconut" and "coffee" were successfully matched, thus assigning the crop type label "rubber tree" to the data record. When assigning geographic region and time period labels, the data labeling module performed natural language parsing on the attribute description information of each data record in the standardized dataset, identifying and extracting proper noun phrases describing geographic location and adverbial phrases describing time period as keywords. For the attribute description information "The rubber forest located in Danzhou City, Hainan Province, was tapped during the rainy season in 2023", the data labeling module extracted the spatial location keyword "Danzhou City, Hainan Province" and the time period keyword "during the rainy season in 2023".

[0023] The data tagging module performs fuzzy matching between extracted spatial location keywords and administrative division names, topographic terms, and climate zone classification names in the geographic climate zoning database. It calculates the semantic similarity between each keyword and the database entry. The geographic climate zoning database contains the entry "Hainan Province, Danzhou City," corresponding to the standard geographic region code "460400." The data tagging module uses a semantic similarity calculation formula for matching, defined as follows:

[0024] in: Indicates semantic similarity. A vocabulary set of keywords indicating spatial location. A vocabulary set representing entries in a geographic and climatic zoning database. This represents the size of the intersection of two sets. and Representing the size of the two sets respectively, when semantic similarity When the matching threshold of 0.8 is exceeded, the data tagging module assigns the standard geographic region code "460400" corresponding to the entry in the geographic climate zoning database to the data record, forming the geographic region tag "460400". It can be understood that the preset matching threshold can be adjusted according to the actual application scenario. The data tagging module aligns the extracted time period keywords with the crop phenological calendar, traditional agricultural solar terms and planting season division information recorded in the historical production season database. The historical production season database records phenological calendars for rubber trees, such as "seedling stage", "growth stage" and "tapping stage". The data tagging module identifies the specific production stage to which the time period keyword "2023 rainy season" belongs in the historical production season database as "tapping stage" and assigns the unique identifier of "tapping stage" to the data record, forming the time stage tag "tapping stage". In some embodiments, the assignment of the time stage tag depends on the completeness of the historical production season database.

[0025] The data tagging module identifies specific agricultural operations or management terms in the attribute description information and compares them with a preset production process list. This list includes terms such as "planting," "fertilizing," "irrigating," "tapping rubber," and "harvesting." The data tagging module identifies "tapping rubber" from the attribute description information and matches it with "tapping rubber" in the production process list, assigning the production process tag "tapping rubber" to the data record. Optionally, the production process list can be expanded to include more agricultural operation terms. Finally, the data tagging module adds the crop type tag "rubber tree," the geographical region tag "460400," and the time tag. The combination of the stage label "rubber tapping period" and the production stage label "rubber tapping" generates a complete classification label for the data record: "Rubber Tree - 460400 - Rubber Tapping Period - Rubber Tapping". In practice, the data tagging module processes another data record containing the entity name "coffee" and the attribute description "Irrigation measures were implemented in Xishuangbanna, Yunnan Province at the beginning of the dry season". The data tagging module extracts the entity name "coffee" and matches it with a tropical crop species name database to assign the crop species label "coffee". It also extracts the spatial location keyword "Xishuangbanna, Yunnan Province" and matches it with a geographic and climatic regionalization database. The entry "Yunnan Province, Xishuangbanna Dai Autonomous Prefecture" corresponds to the code "532800". After calculating semantic similarity, the geographic region label "532800" is assigned. The time period keyword "early dry season" is extracted and aligned with the historical production season database. The historical production season database records "early dry season" for coffee corresponding to "flowering management stage", so the time stage label "flowering management stage" is assigned. The agricultural operation term "irrigation measures" is identified and compared with the production link list, so the production link label "irrigation" is assigned. The complete classification label "coffee-532800-flowering management stage-irrigation" is generated. It can be understood that through this process... In this system, each data record in the standardized dataset is assigned a multi-dimensional classification label. In some embodiments, the data labeling module processes multiple data records in the standardized dataset in batches and compares the label generation process of different data records. For example, the data record "Rubber Tree - Danzhou City, Hainan Province - 2023 Rainy Season - Rubber Tapping" and the data record "Coffee - Xishuangbanna Region, Yunnan - Early Dry Season - Irrigation" differ in crop type label, geographical region label, time stage label, and production link label, demonstrating the distinguishing ability of multi-dimensional labeling processing. Optionally, the combination format of classification labels can be connected by delimiters.

[0026] In one embodiment of the present invention, see [reference] Figure 2The pedigree management module constructs a pedigree framework for tropical agricultural resources. The root node of this framework is defined as the tropical agricultural resource data. Under the root node, the module establishes first-level branch nodes according to major tropical crop categories. These first-level branch nodes can include "Rubber," "Tropical Fruits," "Tropical Beverage Crops," and "Tropical Spice Crops." Under each first-level branch node, the module establishes second-level branch nodes based on different crop varieties or strains. For example, under the first-level branch node "Tropical Beverage Crops," its second-level branch nodes can include "Coffee-Arabica," "Coffee-Robusta," "Cocoa-Forestro," and "Tea-Large Leaf." Under the second-level branch nodes, the module establishes tertiary branch nodes according to different phenological stages within the crop's growth cycle. For example, under the second-level branch node "Coffee-Arabica," its tertiary branch nodes can include... Under the three-level branch nodes—"Seedling stage," "Growth stage," "Flowering stage," "Fruit development stage," and "Harvest stage"—the pedigree management module establishes four-level branch nodes according to different geographical and climatic zones. Taking the three-level branch node "Coffee-Arabica-Flowering stage" as an example, its four-level branch nodes can include "Yunnan dry-hot valley region," "Hainan mountainous and hilly region," and "Guangdong Leizhou Peninsula region." Under the four-level branch nodes, the pedigree management module establishes five-level branch nodes according to the type of resource data. The types of resource data include environmental monitoring data, soil property data, pest and disease occurrence data, agronomic operation record data, and yield and quality data. Taking the four-level branch node "Coffee-Arabica-Flowering stage-Yunnan dry-hot valley region" as an example, its five-level branch nodes are "Environmental monitoring data," "Soil property data," "Pest and disease occurrence data," "Agronomic operation record data," and "Yield and quality data."

[0027] When mapping initial agricultural resource data with classification labels to a tropical agricultural resource phylogenetic framework, the phylogenetic management module reads the classification labels carried by the initial agricultural resource data. Taking initial agricultural resource data with the classification label "Coffee-532800-Flowering Management Stage-Irrigation" as an example, the phylogenetic management module matches the crop type label "Coffee" in the classification label with the names of the first-level and second-level branch nodes in the tropical agricultural resource phylogenetic framework. The first-level branch node "Tropical Beverage Crops" contains "Coffee". Further, under the "Tropical Beverage Crops" node, it matches the second-level branch nodes "Coffee-Arabica" and "Coffee-Robusta". The specific variety needs to be determined based on the data content or additional information. Assuming the target second-level branch node is determined to be "Coffee-Arabica", the phylogenetic management module matches the time stage label "Flowering Management Stage" in the classification label with the name of the third-level branch node under the target second-level branch node "Coffee-Arabica". "Flowering Management Stage" corresponds to the third-level branch node "Flowering Period", thus determining the target third-level branch node to be "Coffee-Arabica-Flowering". The genealogy management module matches the geographic region label "532800" in the classification tags with the name of the fourth-level branch node under the target third-level branch node "Coffee-Arabica-Flowering Period". The geographic region label "532800" corresponds to "Xishuangbanna Dai Autonomous Prefecture, Yunnan Province", which matches the geographic range covered by the fourth-level branch node "Yunnan Dry-Hot Valley Area". Therefore, the target fourth-level branch node is determined to be "Coffee-Arabica-Flowering Period-Yunnan Dry-Hot Valley Area". The genealogy management module matches the theme content of the initial agricultural resource data with the data type description of the fifth-level branch node under the target fourth-level branch node. The initial agricultural resource data describes irrigation measures, which belong to agricultural operations. Therefore, it matches the data type description of "Agronomic Operation Record Data". The final target fifth-level branch node to be stored is "Coffee-Arabica-Flowering Period-Yunnan Dry-Hot Valley Area-Agronomic Operation Record Data", and the initial agricultural resource data is stored therein. In some embodiments, the mapping process uses a node matching weight calculation formula to assist decision-making. The node matching weight calculation formula is defined as:

[0028] in: Indicates the matching weight. The feature vector representing the classification label, The feature vector representing the name of a node in the genealogical framework can be understood as being obtained through word embedding techniques.

[0029] In practice, another initial agricultural resource data with the classification label "rubber tree-460400-tapping season-tapping" undergoes a mapping process. The pedigree management module matches the crop type label "rubber tree" with the first-level branch node "rubber," and further matches second-level branch nodes such as "rubber tree-thermal research 7-33-97 strain" under the "rubber" node to determine the target second-level branch node. The pedigree management module then matches the time stage label "tapping season" with the tertiary branch nodes under the target second-level branch node, such as "tapping season." The matching process identifies the target third-level branch node as "Rubber Tree - Thermal Research 7-33-97 Variety - Tapping Season". The pedigree management module matches the geographic region tag "460400" with the fourth-level branch node under the target third-level branch node, such as "Hainan Danzhou Region", thus identifying the target fourth-level branch node as "Rubber Tree - Thermal Research 7-33-97 Variety - Tapping Season - Hainan Danzhou Region". The pedigree management module then matches the subject content of the initial agricultural resource data, "Tapping Operation Record", with the fifth-level branch node under the target fourth-level branch node. The data is matched with the "Agronomic Operation Record Data" node and ultimately stored in the "Rubber Tree - Tropical Research 7-33-97 Variety - Tapping Period - Hainan Danzhou Area - Agronomic Operation Record Data" node. It can be understood that by comparing the "Coffee - Arabica Variety - Flowering Period - Yunnan Dry-Hot Valley Area - Agronomic Operation Record Data" node with the "Rubber Tree - Tropical Research 7-33-97 Variety - Tapping Period - Hainan Danzhou Area - Agronomic Operation Record Data" node, the data is organized into different branch paths of the phylogenetic framework based on its multidimensional classification labels. In some embodiments, when the information in the classification labels is insufficient to accurately match an existing secondary branch node, optionally, the phylogenetic management module can map the data to a higher-level node or trigger the creation process of a new branch node. For example, for data where the crop type label is only "Coffee" without specifying a variety, it may be temporarily mapped to a general "Coffee" secondary node under the primary branch node "Tropical Beverage Crops". Optionally, the node names and hierarchical structure of the phylogenetic framework can be expanded and modified according to the actual knowledge system.

[0030] In one embodiment of the present invention, see [reference] Figure 3The correlation analysis module establishes cross-data source correlation links within the tropical agricultural resource phylogenetic framework. Under the same target fifth-level branch node, the module filters initial agricultural resource data from different data collection sources. Taking a target fifth-level branch node "Coffee-Arabica-Flowering Period-Yunnan Dry-Hot Valley Area-Environmental Monitoring Data" as an example, this node stores two initial agricultural resource data. The first data comes from an automatic weather station, recording the daily average temperature, relative humidity, and rainfall from March 1st to March 10th, 2023. Its classification labels include the time period label "Flowering Period" and the geographical region label "532800 (Xishuangbanna)". The second data comes from a handheld field sensor, recording data from March 5th to March 7th, 2023. The hourly light intensity and soil temperature data during the daytime also include the time period label "flowering period" and the geographic region label "532800 (Xishuangbanna)" in their classification tags. The correlation analysis module analyzes the coupling relationship between these initial agricultural resource data in terms of time period labels, geographic region labels, and specific data content. In terms of time period labels, both data are labeled "flowering period" and their specific time ranges overlap, which constitutes a temporal continuity relationship. In terms of geographic region labels, both data are labeled "532800" and point to the same geographic climate zone, which constitutes a spatial overlap relationship. In terms of specific data content, the first data provides macro-meteorological elements, while the second data provides more refined micro-environmental elements, and the two form a complementary relationship in describing environmental conditions.

[0031] For initial agricultural resource data exhibiting temporal continuity, spatial overlap, and content complementarity, the association analysis module creates association indexes between them. These indexes record the association attributes between two data sets across time, space, and content dimensions. Based on the type of association index, the module constructs data tracing links, spatial overlay links, or content complementarity links. For the aforementioned two data sets, the module constructs a content complementarity link, indicating that macro-meteorological data from automatic weather stations and micro-environmental data from handheld sensors can be combined in a specific time and space to more comprehensively assess environmental conditions during the flowering period. The module stores this link information in the metadata layer of the tropical agricultural resource genealogy framework. In some embodiments, the module uses an association strength assessment formula to quantify the coupling relationship between data sets. The association strength assessment formula is defined as:

[0032] in: Indicates the strength of the association. Indicates the degree of overlap in time indices. Indicates spatial index overlap. Indicates the complementarity of content indexes. , , These are preset weighting coefficients, which is understandable; only the correlation strength is considered. Only data pairs exceeding a preset threshold will have their association links established.

[0033] The data verification and supplementation module verifies agricultural resource data according to preset data integrity rules. This module defines a data integrity template for each fifth-level branch node in the tropical agricultural resource genealogy framework. Taking the aforementioned fifth-level branch node "Coffee-Arabica-Flowering Period-Yunnan Dry-Hot Valley Area-Environmental Monitoring Data" as an example, its data integrity template ideally includes the data fields "Air Temperature," "Air Humidity," "Rainfall," "Light Intensity," "Soil Temperature," "Soil Humidity," "Wind Speed," and "Wind Direction." The data time span should cover the entire "flowering period," and the data spatial coverage density requires at least one valid data collection point per 100 square kilometers. The data verification and supplementation module retrieves all initially stored agricultural data under the fifth-level branch node "Coffee-Arabica-Flowering Period-Yunnan Dry-Hot Valley Area-Environmental Monitoring Data." The resource data and its established data association links actually store two data points from automatic weather stations and handheld sensors, as well as three other data points from satellite remote sensing, fixed soil moisture stations, and drone patrols, respectively. The data verification and supplementation module compares the field coverage, time span, and spatial density of the actual data with the data integrity template item by item. In terms of field coverage, the actual data set provides "air temperature," "air humidity," "rainfall," "light intensity," and "soil temperature," but lacks the fields of "soil moisture," "wind speed," and "wind direction." In terms of time span, the actual data covers part of the flowering period in 2023, but fails to completely cover the entire period from February to April as specified in the template, resulting in a time gap. In terms of spatial density, the distribution of actual data collection points meets the requirement of at least one point per 100 square kilometers.

[0034] When actual data fails to meet the threshold set by the data integrity template in any aspect—field coverage, time span, or spatial density—the data verification and supplementation module determines that there is a data verification gap in the fifth-level branch node and records the specific type of the gap. In the example above, due to the missing fields of "soil moisture," "wind speed," and "wind direction," and the incomplete time span, the data verification and supplementation module determines that there is a data verification gap in the node "coffee-Arabica variety-flowering period-Yunnan dry-hot valley area-environmental monitoring data" and records the gap type as "key field missing" and "incomplete time series." In specific implementation, comparing the verification process of another fifth-level branch node "rubber tree-Reyan 7-33-97 variety-tapping period-Hainan Danzhou area-yield and quality data," its data integrity template requires the inclusion of fields such as "single tree yield," "dry rubber content," and "impurity content." The time span covers the entire rubber tapping season, and the spatial coverage requirement covers the main planting areas. The data verification and supplementation module retrieves the actual stored data under this node and finds that all fields are complete and the time coverage is complete, but the spatial data for the "Nada Town" area is missing. Therefore, the data verification and supplementation module determines that there is a data verification gap at this node and records the gap type as "incomplete spatial coverage". It is understandable that the threshold setting of the data integrity template can be adjusted according to different data types and application requirements. Optionally, for pest and disease occurrence data, the spatial coverage density requirement may be higher than that for environmental monitoring data. In some embodiments, when comparing the time span, the data verification and supplementation module not only checks the start and end times, but also checks the continuity of the time series. Optionally, when verifying field coverage, the data verification and supplementation module will distinguish between required fields and optional fields, and only the absence of required fields will be judged as a gap.

[0035] In one embodiment of the present invention, the data verification and supplementation module triggers the data supplementation collection process. The data verification and supplementation module generates a structured data supplementation task description according to the specific type of data verification gap. The data supplementation task description specifies the data subject, target time range, target geographical area and required data fields to be supplemented. Taking a fifth-level branch node "coffee-Arabica-flowering period-Yunnan dry-hot valley area-environmental monitoring data" with gaps of "missing key fields" and "incomplete time series" as an example, refer to Table 1 for an exemplary data supplementation task description generated by the data verification and supplementation module.

[0036] Table 1: Description of Data Supplementation Task Task elements Specific content Data topic Environmental monitoring data Target time range February 15, 2023 to April 15, 2023 Target geographic area The dry-hot valley region under the geographic code 532800 (Xishuangbanna Dai Autonomous Prefecture, Yunnan Province) Required data fields Soil moisture, wind speed, wind direction Gap type Key fields missing, time series incomplete The data verification and supplementation module publishes the data supplementation task description to the preset data acquisition task queue. The data acquisition task queue automatically matches the device or data source interface with corresponding data acquisition capabilities according to the task description. For example, the data acquisition task queue matches the above task with an IoT weather station node deployed in Xishuangbanna that has the ability to monitor soil moisture, wind speed and wind direction. The data verification and supplementation module sends a data acquisition command to the successfully matched IoT weather station node and receives the returned supplementary acquisition data. The supplementary acquisition data may include the daily average soil moisture, maximum wind speed and wind direction frequency data from February 20 to April 10, 2023. The data verification and supplementation module performs format standardization and multi-dimensional labeling processing on the supplementary acquisition data to form new initial agricultural resource data. The new initial agricultural resource data is assigned the classification label "coffee-532800-flowering period-environmental monitoring" and mapped to fill the original data verification gap in the fifth-level branch node "coffee-Arabica-flowering period-Yunnan dry and hot valley area-environmental monitoring data".

[0037] After mapping and filling in the new initial agricultural resource data, the data verification and supplementation module performs a correlation link update step. This module checks the classification labels and specific content of the new initial agricultural resource data. The classification label for the new initial agricultural resource data is "Coffee-532800-Flowering Period-Environmental Monitoring," and the specific content includes soil moisture, wind speed, and wind direction data. Under the same five-level branch node "Coffee-Arabica-Flowering Period-Yunnan Dry-Hot Valley Area-Environmental Monitoring Data," the data verification and supplementation module automatically calculates the correlation between the new initial agricultural resource data and the existing initial agricultural resource data. The data verification and supplementation module uses a correlation calculation equation for evaluation, defined as follows:

[0038] in: The score represents the correlation calculation. Indicates the time overlap coefficient. Indicates the spatial overlap coefficient. Indicates the content complementarity coefficient. , , It is a preset normalized weighting factor and a time overlap coefficient. Based on the temporal overlap between new initial agricultural resource data and existing data, the spatial overlap coefficient is calculated. Consistency calculation based on geographic region tags, content complementarity coefficient Based on the complementarity calculation of data fields, if the correlation calculation score is... If the threshold is exceeded, a new association is determined to have been discovered.

[0039] If a new correlation is discovered, the data verification and supplementation module establishes a new data association link between the new initial agricultural resource data and related existing initial agricultural resource data, or extends and expands the existing data association link. For example, if new soil moisture, wind speed, and wind direction data completely overlap with existing automatic weather station temperature, humidity, and rainfall data within a node in time and space, and their fields are complementary, the data verification and supplementation module therefore establishes a new data content complementary link to associate these three types of data. The data verification and supplementation module updates the link information stored in the metadata layer of the tropical agricultural resource genealogy framework and stores the index information of the newly established data content complementary link in the metadata layer. In some embodiments, the weighting factor in the correlation calculation equation... , , Different configurations can be made based on different data types. It is understandable that the setting of preset thresholds affects the sensitivity of establishing associated links.

[0040] In practical implementation, compared with another data supplementation and correlation link update scenario, the data verification and supplementation module generates a data supplementation task description for a five-level branch node with a "spatial incomplete coverage" gap: "Rubber Tree - Thermal Research 7-33-97 Variety - Tapping Period - Hainan Danzhou Area - Yield and Quality Data". The data supplementation task description requires supplementing yield data for the "Nada Town" area. The data collection task queue matches a data entry terminal that manually records data in that area. The data verification and supplementation module receives the supplemented yield data, processes it to form new initial agricultural resource data, and maps and fills the gap. In the link update step, the data verification and supplementation module calculates the correlation between the new yield data and the existing yield data of other regions. Since they are spatially adjacent and the crop varieties and management are consistent, the data verification and supplementation module establishes a data spatial overlay link for cross-regional yield comparison analysis. Optionally, for time series data, the link update step may include extending the original data tracing link. In some embodiments, the data supplementation task description can also specify the frequency or accuracy requirements of data collection. Optionally, the matching logic of the data collection task queue can perform similarity calculation based on the device capability description library.

[0041] See Figure 4This is a bar chart comparing the completeness of coffee environmental monitoring data before and after supplementary data collection. It primarily shows the changes in the completeness of five monitoring fields before and after supplementary data collection. The initial completeness of "temperature" and "rainfall" was already high (75%-80%), and only slightly improved after supplementary collection, indicating that these two fields already had a good data foundation. This chart visually verifies the actual effect of the "data verification and supplementation module," corresponding to the supplementary data collection scenario of the "coffee-Arabica variety-flowering period-Yunnan dry-hot valley area-environmental monitoring data" node in your project. It clearly demonstrates that the supplementary data collection task accurately located the problems of "missing key fields" and "incomplete time series." The data completeness after supplementary collection meets business requirements, providing a reliable foundation for subsequent data correlation analysis and decision-making.

[0042] In one embodiment of the present invention, the system performs a data credibility pre-screening step before standardizing the data format of the original data records. This pre-screening step assigns a source credibility score to each original data record. The source credibility score is calculated based on the historical data quality rating of the data provider, the calibration status of the data acquisition equipment, and the integrity of the data transmission link. The calculation of the source credibility score follows a preset quantitative evaluation model. The quantitative evaluation model is as follows:

[0043] in: This represents the source credibility score, indicating the historical data quality rating of the data provider. This represents the calibration status coefficient of the data acquisition equipment. This represents the integrity coefficient of the data transmission link. , , These are preset weighting factors corresponding to the three indicators, and Historical data quality rating The calibration status coefficient is determined based on the accuracy and completeness of previously provided data from the data provider after verification. The integrity coefficient of the data transmission link is 1.0 when the device has recently undergone standard calibration, 0.5 when the calibration is expired or the status is unknown, and 0.1 when there is no calibration record. The score is determined based on whether the data transmission log is complete and whether there are any signs of interruption or tampering. A score of 1.0 is given when the log is complete and without any abnormalities, 0.7 is given when there is a short-term interruption that can be repaired, and 0.3 is given when there are abnormalities or missing data.

[0044] The data credibility pre-screening step sets a credibility threshold to filter out raw data records whose source credibility scores are lower than the threshold. The system only performs subsequent format standardization processing on the raw data records that pass the credibility screening. In a specific example, the system receives two raw data records about soil moisture monitoring of the same plot. The first raw data record comes from an automatic weather station that has undergone annual calibration and has a complete transmission log. Its data provider's historical data quality rating is 0.9, calibration status coefficient is 1.0, and transmission link integrity coefficient is 1.0. Assuming a weighting factor... ,, Then calculate its source credibility score. The second raw data record comes from a simple sensor that has no recorded calibration history and is transmitted over an unstable mobile network. Its data provider has a historical data quality rating of 0.6, a calibration status coefficient of 0.1, and a transmission link integrity coefficient of 0.3. Calculate its source reliability score. Assuming the system's credibility threshold is set to 0.7, the first original data record with a source credibility score of 0.95 will pass the screening and enter the subsequent processing flow, while the second original data record with a source credibility score of 0.39 will be filtered out.

[0045] In some embodiments, historical data quality rating values The calculation can be further refined, for example, by introducing a time decay factor to give greater weight to the impact of recent data quality on the rating. After calculating the source credibility score, the data credibility pre-screening step can record the score details for subsequent audit traceability. It can be understood that the credibility threshold can be dynamically adjusted according to the rigor of the data application scenario. For high-precision agricultural model building scenarios, the credibility threshold can be set to 0.8; for general trend analysis scenarios, the credibility threshold can be set to 0.6.

[0046] In another comparative example, the system processes a field observation record from a well-known agricultural research institution and a text report from an anonymous online source. The well-known agricultural research institution has a historical data quality rating of 0.95, its observation equipment calibration status coefficient is 1.0, the data is transmitted via an encrypted dedicated line, and its integrity coefficient is 1.0, resulting in a source credibility score higher than 0.9. The anonymous online source has no historical data quality rating record, therefore... The default value of 0.3 is used. Since there is no device information, the calibration status coefficient is set. The integrity coefficient is set to 0.1 because the transmission link is ambiguous. Taking a value of 0.3, if the source credibility score is lower than 0.3, and with a credibility threshold of 0.7, the former passes the screening, while the latter is filtered. In some embodiments, the data credibility pre-screening step marks and archives the filtered original data records instead of directly and permanently deleting them. Optional, a weighting factor... , , The values ​​can be configured by the system administrator according to different data types. For example, for real-time sensor data, the focus may be more on the device calibration status and transmission integrity, while for manually entered statistical reports, the focus may be more on the historical rating of the data provider.

[0047] See Figure 5 This is a data verification and supplementation effectiveness evaluation chart, showing the changes in the completeness of five core data types in the tropical agricultural resource database before and after supplementation, as well as the corresponding data supplementation rates. The completeness of all data types improved to over 85% after supplementation, indicating that the data verification and supplementation module effectively filled various data gaps. This chart clearly reflects the effectiveness of the "data verification and supplementation module" on different data types. High supplementation rates correspond to typical gaps in business operations such as "missing key fields" and "incomplete time series," and the improved data completeness after supplementation directly supports subsequent correlation analysis and decision-making. Low supplementation rates indicate that the original data collection mechanism for this type of data is more complete, which can provide a reference for optimizing the collection process of other data.

[0048] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.

Claims

1. A system for constructing a tropical agricultural resource database based on agricultural big data, characterized in that: include: The data preprocessing module collects various raw data records from the field of tropical agriculture, performs data format standardization processing on the raw data records, and forms a standardized data set. The data tagging module performs multi-dimensional tagging on standardized datasets based on tropical crop types, geographical and climatic zones, and production management processes, generating initial agricultural resource data with classification tags. The genealogy management module constructs a tropical agricultural resource genealogy framework, mapping initial agricultural resource data with classification labels to the corresponding hierarchical nodes of the tropical agricultural resource genealogy framework according to their classification labels. The correlation analysis module performs data correlation analysis on initial agricultural resource data at the same level node within the tropical agricultural resource genealogy framework, and establishes data correlation links across data sources. The data verification and supplementation module verifies agricultural resource data organized through data association links according to preset data integrity rules. Data that fails verification is marked as data verification gaps, and the data supplementation and collection process is triggered.

2. The tropical agricultural resource database construction system based on agricultural big data according to claim 1, characterized in that, The multidimensional labeling of standardized datasets based on tropical crop types, geographical and climatic zones, and production management processes includes the following steps: Parse each data record in the standardized dataset and extract the entity object name and attribute description information contained in the data record; Match the entity object name with the tropical crop species name database and assign crop species labels to the data records; Analyze the spatial location keywords and time period keywords in the attribute description information, compare them with the geographic climate zoning database and the historical production season database respectively, and assign geographic area labels and time period labels to the data records. Identify specific agricultural operations or management measures terms involved in the attribute description information, compare them with the preset production link list, and assign production link labels to the data records; By combining crop type labels, geographical region labels, time period labels, and production stage labels, classification labels corresponding to data records are generated.

3. The tropical agricultural resource database construction system based on agricultural big data according to claim 2, characterized in that, Constructing a phylogenetic framework for tropical agricultural resources includes the following steps: The root node of the phylogenetic framework is defined as tropical agricultural resource data; Under the root node, establish first-level branch nodes according to the major tropical crop categories; Under each primary branch node, secondary branch nodes are established based on different crop varieties or strains; Under the secondary branch nodes, tertiary branch nodes are established according to the different phenological stages within the crop growth cycle; Under the third-level branch nodes, fourth-level branch nodes are established according to different geographical and climatic zones; Under the fourth-level branch node, fifth-level branch nodes are established according to the types of resource data. The types of resource data include environmental monitoring data, soil property data, pest and disease occurrence data, agronomic operation record data, and yield and quality data.

4. The tropical agricultural resource database construction system based on agricultural big data according to claim 3, characterized in that, Mapping initial agricultural resource data with classification labels to corresponding hierarchical nodes in the tropical agricultural resource genealogy framework according to their classification labels includes the following steps: Read the classification labels carried by the initial agricultural resource data; Match the crop type labels in the classification tags with the names of the first-level and second-level branch nodes in the tropical agricultural resource genealogy framework to determine the target second-level branch node; Match the time period label in the category label with the name of the third-level branch node under the target second-level branch node to determine the target third-level branch node; Match the geographic region label in the category label with the name of the fourth-level branch node under the target third-level branch node to determine the target fourth-level branch node; The subject content of the initial agricultural resource data is matched with the data type description of the fifth-level branch node under the target fourth-level branch node to determine the final target fifth-level branch node for storage, and the initial agricultural resource data is stored in it.

5. The tropical agricultural resource database construction system based on agricultural big data according to claim 4, characterized in that, Establishing a data link across data sources involves the following steps: Under the same target level 5 branch node, initial agricultural resource data from different data collection sources are filtered out; Analyze the coupling relationship between initial agricultural resource data in terms of time period labels, geographic region labels, and specific data content; For initial agricultural resource data that have temporal continuity, spatial overlap, or complementary content, create association indexes between them; Based on the type of associated index, construct data traceability links, data spatial overlay links, or data content complementary links, and store the link information in the metadata layer of the tropical agricultural resource genealogy framework.

6. The tropical agricultural resource database construction system based on agricultural big data according to claim 5, characterized in that, Verifying agricultural resource data organized via data association links according to preset data integrity rules includes the following steps: A data integrity template is defined for each fifth-level branch node in the tropical agricultural resource genealogy framework. The data integrity template specifies the data fields, data time span, and data spatial coverage density that should be included under the ideal state of the fifth-level branch node. Retrieve all the initial agricultural resource data actually stored under a certain fifth-level branch node and the data association links established thereunder; The actual data's field coverage, time span, and spatial density are compared item by item with the data integrity template. When the actual data fails to meet the threshold set by the data integrity template in any aspect such as field coverage, time span, or spatial density, it is determined that there is a data verification gap in the fifth-level branch node, and the specific type of the gap is recorded.

7. The tropical agricultural resource database construction system based on agricultural big data according to claim 6, characterized in that, The process of triggering supplementary data collection includes the following steps: Based on the specific type of data verification gap, a structured data supplementation task description is generated. The data supplementation task description clearly specifies the data subject to be supplemented, the target time range, the target geographical area, and the required data fields. Publish the data supplementation task description to the preset data acquisition task queue; The data acquisition task queue automatically matches devices or data source interfaces with corresponding data acquisition capabilities based on the task description; Send data acquisition instructions to the successfully matched device or interface, and receive the supplementary acquisition data returned. The supplementary collected data is formatted and multidimensionally labeled to form new initial agricultural resource data, which is then mapped to fill the original five-level branch nodes where data verification gaps existed.

8. The tropical agricultural resource database construction system based on agricultural big data according to claim 7, characterized in that, After mapping and filling the gaps in the original data verification of the fifth-level branch nodes, the process also includes an associated link update step: Examine the classification labels and specific details of the new initial agricultural resource data; Under the same fifth-level branch node, the correlation between new initial agricultural resource data and existing initial agricultural resource data is automatically calculated; If new relationships are discovered, a new data link is established between the new initial agricultural resource data and the relevant existing initial agricultural resource data, or the existing data link is extended and expanded. Update the link information stored in the metadata layer of the tropical agricultural resource genealogy framework.

9. The tropical agricultural resource database construction system based on agricultural big data according to claim 1, characterized in that, Before standardizing the data format of the original data records, a data credibility pre-screening step is also included: Each original data record is assigned a source credibility score, which is calculated based on the data provider's historical data quality rating, the calibration status of the data acquisition equipment, and the integrity of the data transmission link. Set a credibility threshold to filter out original data records whose source credibility score is lower than the credibility threshold; Subsequent format standardization processing is performed only on the original data records that have passed the credibility screening.

10. The tropical agricultural resource database construction system based on agricultural big data according to claim 2, characterized in that, The spatial location keywords and time period keywords in the analytical attribute description information are compared with the geographic climate zoning database and the historical production season database, respectively, and the data records are assigned geographic area labels and time period labels, including the following steps: Natural language parsing is performed on the attribute description information of each data record in the standardized dataset to identify and extract proper noun phrases describing geographical location and adverbial phrases describing time period as keywords; The extracted spatial location keywords are fuzzily matched with the administrative division names, topographic terms and climate zone classification names in the geographic climate zoning database, and the semantic similarity between each keyword and the database entry is calculated. When the semantic similarity exceeds the preset matching threshold, the standard geographic region code corresponding to the entry in the geographic climate zoning database is assigned to the data record to form a geographic region label; The extracted time period keywords are aligned with the crop phenological calendar, traditional agricultural solar terms, and planting season division information recorded in the historical production season database. Identify the specific production stage in the historical production season database to which the time period indicated by the time cycle keyword belongs, and assign a unique identifier to the data record for the specific production stage to form a time stage label.