Agricultural knowledge graph construction method based on large model
By constructing an agricultural knowledge graph using a large model, the problem of handling semantic relationships in agricultural policies in existing technologies has been solved. This enables efficient policy conflict detection and compliance monitoring, improves the accuracy and timeliness of the agricultural knowledge graph, and meets the real-time compliance requirements of precision agriculture.
Patent Information
- Application Number
- CN202510962688.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-10-28
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing agricultural knowledge graph technologies struggle to handle the complex semantic relationships and spatiotemporal dynamics of agricultural policies, resulting in limitations in cross-regional policy adaptability, implicit conflict identification, and decision interpretability. They fail to meet the real-time compliance supervision requirements of precision agriculture. Furthermore, traditional negative sampling methods generate pseudo-negative samples that violate agricultural norms, causing knowledge graphs to fail in compliance-sensitive scenarios.
This paper adopts a large-model-based agricultural knowledge graph construction method. Structured policy texts are acquired through a distributed data acquisition system, and deep semantic features are extracted by a bidirectional LSTM network with an attention mechanism. A semantic association graph of policy clauses is constructed, and the semantic distance between nodes is calculated using a graph neural network. The conflict judgment threshold is dynamically adjusted by combining a timeliness weight factor to generate quantitative policy conflict feature values. The model is validated through a multimodal strategy for risk assessment and optimization, and an adaptive optimization strategy is implemented to generate a compliance-type knowledge graph.
It has achieved accurate analysis of policy texts and real-time compliance monitoring, improved the accuracy of policy conflict detection to 92.3%, shortened the time to discover pesticide use violations to 2.3 days, improved the timeliness and interpretability of the system, and met the needs of dynamic updates and cross-regional adaptation of agricultural policies.
Smart Images

Figure CN120851163A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of agricultural informatization and digital agriculture technology, specifically to a method for constructing agricultural knowledge graphs based on large models. Background Technology
[0002] With the accelerated digital transformation of agriculture, policy compliance management faces the challenge of handling massive, multi-source, and dynamically updated agricultural regulatory texts. Traditional manual review methods are inefficient and prone to oversights. Existing agricultural knowledge graph technologies are mostly based on rule engines or shallow machine learning, which struggle to handle the complex semantic relationships and spatiotemporal dynamics of policy texts. Although some research has attempted to introduce deep learning, it still has significant limitations in areas such as cross-regional policy adaptability, implicit conflict identification, and decision interpretability, failing to meet the real-time compliance supervision requirements of precision agriculture. The frequent updates to agricultural policies (an average annual revision rate of 23%) and regional differences (standard differences between provinces and cities reaching 35%) further exacerbate the challenges of knowledge graph construction and maintenance.
[0003] In the training process of Knowledge Graph Embedding (KGE) models, traditional negative sampling methods generate negative samples by randomly replacing the first and last entities of triples, but without introducing domain knowledge constraints, resulting in the generation of pseudo-negative samples that violate agricultural norms or policies. This domain-agnostic sampling mechanism causes the model to implicitly learn incorrect domain logic. When the decoder makes relation predictions based on the contaminated vector space, it may output harmful recommendations that violate regulatory requirements (such as recommending banned pesticides in organic farming scenarios). The core of this problem lies in the lack of a domain rule injection module and a policy-based negative sample filtering layer in existing sampling algorithms, which prevents the model from distinguishing between "randomly unreasonable" and "domain-illegal" negative samples, ultimately causing the knowledge graph to fail in agricultural compliance-sensitive scenarios. It is necessary to impose hard constraints on the sampling space through a structured rule engine (such as a pesticide use knowledge subgraph based on REACH regulations) and design a domain-adaptive adversarial negative sampling strategy to eliminate the risk of toxic recommendations. Summary of the Invention
[0004] The purpose of this invention is to provide a method for constructing agricultural knowledge graphs based on large models, so as to solve the problems mentioned above.
[0005] The objective of this invention can be achieved through the following technical solutions:
[0006] The method for constructing agricultural knowledge graphs based on large models includes the following steps:
[0007] S1: Construct a distributed data acquisition system to acquire structured policy texts, numerical threshold standards, and logical constraint rules in real time. Dynamically capture updates of multi-source agricultural policy data through adaptive crawling technology, and adopt a data quality verification mechanism to ensure the integrity and consistency of input data.
[0008] S2: A bidirectional LSTM network based on an attention mechanism is used to extract deep semantic features of policy texts, construct a semantic association graph of policy clauses, use a graph neural network to calculate the semantic distance between nodes, and dynamically adjust the conflict judgment threshold by combining the timeliness weight factor to generate quantitative policy conflict feature values.
[0009] S3: Establish a Gaussian mixture model to analyze the threshold distribution characteristics, use a dynamic sliding window algorithm to detect abnormal data points, develop a time series prediction model to warn of threshold drift trends, and generate compliance deviation characteristic values of pesticide combination characteristics, regional differences, and time dimensions for integrated crops.
[0010] S4: The policy conflict feature value and the compliance deviation feature value are integrated into a comprehensive risk feature vector, which is input into a multimodal strategy verification model based on deep residual network. The meta-learning method is used to conduct risk assessment in combination with historical compliance data, and risk source analysis is provided through interpretable AI components.
[0011] S5: Implement adaptive optimization strategies based on risk assessment results, including enhancing the verification strength of high-risk nodes, optimizing negative sampling constraints, and dynamically scheduling computing resources. Through iterative optimization via a closed-loop feedback mechanism, a policy-compliant agricultural knowledge graph that supports compliance monitoring and risk early warning is ultimately generated.
[0012] As a further aspect of the present invention: the process for obtaining the policy conflict feature value is as follows:
[0013] A hierarchical semantic parsing framework is used to process policy texts. First, basic semantic features are extracted by word-level bidirectional LSTM. Then, paragraph-level attention mechanism is used to focus on key clauses. Finally, document-level graph neural network is used to model the global relationship between policy clauses.
[0014] Construct a dynamically updated policy semantic association graph, where nodes represent policy clauses, and edge weights are jointly determined by semantic similarity and timeliness similarity. Semantic similarity is calculated based on deep semantic matching of clause content, while timeliness similarity considers the time interval between policy releases and revision history.
[0015] A multi-dimensional conflict detection algorithm is designed to analyze the semantic contradictions, timeliness conflicts, and differences in authority levels of policy provisions, and output a weighted conflict score matrix.
[0016] As a further aspect of the present invention: the hierarchical semantic parsing specifically includes:
[0017] In the word-level feature extraction stage, a bidirectional LSTM network is used to capture the contextual semantic information of the policy text;
[0018] The paragraph-level processing stage uses a multi-head attention mechanism to identify key constraints in the clauses;
[0019] In the document-level integration phase, semantic relationships across clauses are established using graph attention networks;
[0020] The final policy text representation combines local details with global structure.
[0021] As a further aspect of the present invention: the multi-dimensional conflict detection specifically includes:
[0022] Semantic contradiction detection: Analyze whether there are direct conflicts in the logical constraints between clauses;
[0023] Timeliness conflict assessment: Determine the substitution relationship and transitional provisions between the old and new policies;
[0024] Authority level verification: Different conflict determination standards are set for policies at different levels;
[0025] Comprehensive score generation: The final conflict feature value is output by weighted integration of the detection results from various dimensions.
[0026] As a further aspect of the present invention: the process for obtaining the compliance deviation feature value is as follows:
[0027] Construct a threshold feature library for crop and pesticide combinations, establish independent threshold distribution models for different crop types and pesticide varieties, and analyze the standard distribution patterns of usage in historical data through machine learning algorithms to identify typical threshold features of various combinations;
[0028] A dynamic anomaly detection mechanism is implemented, which automatically adjusts the detection window size according to the crop growth cycle stage, uses multi-dimensional outlier identification technology to find abnormal data points, and combines it with an expert knowledge base for manual review and confirmation.
[0029] Develop a threshold trend prediction system that integrates time series analysis and deep learning technologies to establish a prediction model that takes into account factors such as seasonal changes, regional differences, and policy adjustments, and provide early warnings of potential compliance risks.
[0030] As a further aspect of the present invention: the threshold feature library specifically includes:
[0031] Establish a classification system based on crop type and pesticide type;
[0032] Collect historical usage data to build a distribution model;
[0033] Mark typical threshold ranges and abnormal features;
[0034] Configure an automatic model update mechanism.
[0035] As a further aspect of the present invention: the warning of potential compliance risks specifically includes:
[0036] The detection period is divided according to the crop growth cycle. Multi-scale sliding window technology is used, combined with rule engine and machine learning to identify anomalies and establish an anomaly data hierarchical processing flow.
[0037] Historical threshold change data are collected, seasonal and regional characteristics are analyzed, and a deep learning prediction model is constructed. Based on the output of the deep learning prediction model, a multi-level early warning mechanism is set up. The first level is to quickly screen for threshold overruns based on the statistical model. The second level is to identify abnormal patterns by combining spatiotemporal characteristics. The third level introduces an expert knowledge base for in-depth verification. The fourth level conducts cross-regional risk correlation analysis.
[0038] As a further aspect of the present invention, the construction of the multimodal strategy verification model includes the following steps:
[0039] The design incorporates a feature fusion layer that dynamically fuses policy conflict feature values and compliance deviation feature values using an attention-weighted mechanism.
[0040] Assign semantic weights to text features;
[0041] Assign confidence weights to numerical features;
[0042] Assign timeliness weights to spatiotemporal features;
[0043] We construct the main architecture of a deep residual network, adopt cross-layer connection to avoid gradient vanishing, and embed a domain knowledge-guided convolutional kernel initialization strategy in each residual block.
[0044] Develop a meta-learning training framework to extract meta-features from historical compliance cases and build a context-aware model parameter rapid adaptation mechanism;
[0045] Integrating interpretable AI components enables traceability of the risk assessment process through feature importance analysis and decision path visualization.
[0046] As a further aspect of the present invention: the meta-learning training framework specifically includes:
[0047] Build a multi-granularity compliance case library, marking the key features and outcomes of each case;
[0048] Design a scenario encoder to extract meta-feature representations of cases;
[0049] Develop a parameter prediction network to dynamically generate model parameters based on the meta-features of new cases;
[0050] Establish an online update mechanism to continuously optimize the meta-knowledge representation.
[0051] As a further aspect of the present invention: the implementation of the adaptive optimization strategy specifically includes:
[0052] Establish a hierarchical resource scheduling mechanism to dynamically allocate verification resources based on node risk scores, specifically including:
[0053] If a node's risk score is greater than or equal to the first preset threshold, it is recorded as a high-risk node, and a three-level verification process is initiated, including machine verification, expert review, and cross-system comparison.
[0054] If a node's risk score is less than the first preset threshold but greater than or equal to the second preset threshold, it is designated as a medium-risk node, and a combination of machine verification and sampling review is used.
[0055] If the node risk score is less than the second preset threshold, it is recorded as a low-risk node and a regular inspection mechanism is implemented.
[0056] Develop an intelligent negative sampling optimization system that automatically adjusts the sampling strategy based on risk characteristics, increases sampling density in high-risk areas, sets special sampling constraints for policy-sensitive relationships, introduces adversarial sample generation technology to improve model robustness, and establishes a continuous evaluation mechanism for sampling effects.
[0057] Implement a flexible resource allocation scheme, build a dynamically scalable computing resource pool, implement priority scheduling according to the criticality of tasks, develop a parallel processing framework that supports load balancing, and monitor resource utilization efficiency indicators in real time.
[0058] A self-optimizing closed-loop system is constructed to continuously collect data on the effectiveness of strategy execution, deeply analyze the correlation between strategy parameters and optimization results, automatically adjust the configuration parameters of the optimization strategy, and generate new training datasets through effect verification to drive the continuous evolution of the system.
[0059] The beneficial effects of this invention are:
[0060] (1) This invention innovatively integrates multimodal feature analysis and deep learning technologies to construct an intelligent and precise agricultural policy compliance analysis system. At the technical implementation level, the system adopts a hierarchical semantic parsing architecture. Through a three-level processing flow—word-level bidirectional LSTM network (256-dimensional hidden layers), paragraph-level multi-head attention mechanism (8 heads), and document-level graph attention network (2-layer GAT)—it achieves comprehensive parsing of policy texts from micro-level semantics to macro-level correlations. Test data shows that this architecture achieves an accuracy of 92.3% in policy conflict detection tasks, an improvement of 22.5 percentage points compared to traditional rule-based methods and 15.6 percentage points compared to a single neural network model. Regarding timeliness, the system, through the synergistic effect of a dynamic sliding window algorithm (window size adaptively adjusted from 7 to 30 days) and a spatiotemporal attention LSTM prediction model (inputting 28-dimensional features and outputting 30-day predictions), significantly reduces the average discovery time of pesticide use violations from 17.6 days of traditional manual review to 2.3 days, improving timeliness by 86.9%. This breakthrough is mainly attributed to three core technological advancements: First, the hierarchical semantic parsing architecture, enhanced with pre-trained word vectors (300-dimensional) in the agricultural domain and policy metadata, achieves accurate capture of the deep semantics of policy texts, with an F1 score of 0.91. Second, the dynamically updated threshold analysis model, employing a Gaussian mixture model (3-5 components) and the isolated forest algorithm, can track changes in compliance standards in real time, with daily incremental updates ensuring model timeliness. Third, the multimodal fusion verification mechanism, through the organic combination of a deep residual network (50 layers) and three-channel attention weighting (text 0.5, numerical 0.3, spatiotemporal 0.2), achieves collaborative analysis of textual, numerical, and spatiotemporal features, with an AUC of 0.963.
[0061] (2) This invention achieves adaptive evolution and credible decision-making of agricultural knowledge graphs by constructing a dual-driven intelligent governance system of "meta-learning + interpretable A". In terms of technical implementation, the system first establishes a meta-knowledge base containing more than 100,000 labeled cases, uses a graph neural network architecture to extract 256-dimensional meta-features from the scenario encoder, and dynamically generates model parameters through a hypernetwork, which improves the system's adaptation speed to new policy scenarios by 5 times and maintains a stable accuracy of 94.2% (standard deviation <1.8%) in cross-provincial and municipal tests. The interpretability system integrates four core modules: feature importance analysis based on ensemble gradient (quantifying the contribution of 37 features), GNNExplainer-driven decision path visualization (generating evidence chains with confidence), counterfactual case comparison analysis (identifying the path of least change), and structured audit report generation (supporting PDF / JSON format output), which makes the model transparency SHAP value reach 0.87, meeting the transparency requirements of the EU AI Act. The self-optimization mechanism adopts a dual-loop design: the inner loop fine-tunes parameters based on real-time collected verification accuracy (3000+ samples per day) and processing latency (controlled within 300ms); the outer loop performs a global model update monthly, and triggers knowledge reconstruction in conjunction with the policy change detection module (accuracy of 98.5%), achieving an average monthly performance improvement of 3.2%. Attached Figure Description
[0062] The invention will now be further described with reference to the accompanying drawings.
[0063] Figure 1 This is a flowchart of the agricultural knowledge graph construction method based on a large model according to the present invention. Detailed Implementation
[0064] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.
[0065] Please see Figure 1 As shown, this invention is a method for constructing agricultural knowledge graphs based on large models, including the following steps:
[0066] S1: Construct a distributed data acquisition system to acquire structured policy texts, numerical threshold standards, and logical constraint rules in real time. Dynamically capture updates of multi-source agricultural policy data through adaptive crawling technology, and adopt a data quality verification mechanism to ensure the integrity and consistency of input data.
[0067] S2: A bidirectional LSTM network based on an attention mechanism is used to extract deep semantic features of policy texts, construct a semantic association graph of policy clauses, use a graph neural network to calculate the semantic distance between nodes, and dynamically adjust the conflict judgment threshold by combining the timeliness weight factor to generate quantitative policy conflict feature values.
[0068] S3: Establish a Gaussian mixture model to analyze the threshold distribution characteristics, use a dynamic sliding window algorithm to detect abnormal data points, develop a time series prediction model to warn of threshold drift trends, and generate compliance deviation characteristic values of pesticide combination characteristics, regional differences, and time dimensions for integrated crops.
[0069] S4: The policy conflict feature value and the compliance deviation feature value are integrated into a comprehensive risk feature vector, which is input into a multimodal strategy verification model based on deep residual network. The meta-learning method is used to conduct risk assessment in combination with historical compliance data, and risk source analysis is provided through interpretable AI components.
[0070] S5: Implement adaptive optimization strategies based on risk assessment results, including enhancing the verification strength of high-risk nodes, optimizing negative sampling constraints, and dynamically scheduling computing resources. Through iterative optimization via a closed-loop feedback mechanism, a policy-compliant agricultural knowledge graph that supports compliance monitoring and risk early warning is ultimately generated.
[0071] In S1, a distributed data acquisition system is built to acquire structured policy texts, numerical threshold standards, and logical constraint rules in real time. Adaptive web crawling technology is used to dynamically capture updates of multi-source agricultural policy data, and a data quality verification mechanism is employed to ensure the integrity and consistency of the input data. Specifically, this includes:
[0072] In the data acquisition and preprocessing stage, this invention constructs an intelligent distributed data acquisition system, specifically implemented as follows: The system adopts a microservice architecture to deploy multiple data acquisition nodes, establishing real-time connections with authoritative data sources such as agricultural government departments and standardization organizations through API interfaces. For structured policy texts, the system is configured with a deep learning-based intelligent parser, capable of automatically identifying and extracting key elements such as clauses and constraints from policy documents; for numerical threshold standards, a dedicated tabular data extraction module has been developed, supporting automatic conversion and normalization of various formats such as PDF and Excel; in terms of logical constraint rule acquisition, semantic parsing technology is used to transform rules described in natural language into machine-processable logical expressions. The system integrates an adaptive crawler engine, capable of dynamically adjusting the acquisition frequency and strategy according to the characteristics of different data sources, and automatically triggering incremental acquisition processes when policy updates are detected. To ensure data quality, the system implements a multi-layered verification mechanism: checking data format standardization at the syntactic level, verifying data logical consistency at the semantic level, and reviewing data compliance at the business level, and establishing an automatic repair and manual review process for abnormal data. All collected data will undergo unified data cleaning and standardization processing to ultimately output high-quality data that meets the requirements for knowledge graph construction.
[0073] In S2, a bidirectional LSTM network based on an attention mechanism is used to extract deep semantic features of policy texts, construct a semantic association graph of policy clauses, calculate the semantic distance between nodes using a graph neural network, and dynamically adjust the conflict determination threshold by combining a timeliness weight factor to generate quantitative policy conflict feature values, specifically including:
[0074] In the hierarchical semantic parsing module, the system employs a three-level processing flow for structured analysis of policy text. First, in the word-level feature extraction stage, a bidirectional LSTM network architecture is used to process the segmented policy text, with a 256-dimensional hidden layer and 300-dimensional pre-trained word vectors for the agricultural domain. This captures the context-related semantic representations of words through bidirectional information flow. Second, in the paragraph-level processing stage, an 8-head attention mechanism is configured for in-depth analysis of policy clauses, with a maximum processing length of 512 words, and positional bias is introduced to enhance the identification of key constraints. Finally, in the document-level integration stage, a fully connected policy clause relationship graph is constructed, employing a 2-layer graph attention network for information propagation, with a neighbor sampling count of 8, ultimately generating a text representation that integrates local details and global structure.
[0075] The dynamic semantic association graph construction module achieves association modeling of policy clauses through multi-source feature fusion. The node feature generation stage performs dimensionality reduction and fusion of word-level, paragraph-level, and document-level features, outputting a 128-dimensional feature vector and adding policy metadata. Edge weight calculation employs a weighted combination of semantic similarity and timeliness similarity. Semantic similarity is measured using cosine similarity, while timeliness similarity uses an exponential decay function considering the publication time interval. A semantic weight of 0.6 and a timeliness weight of 0.4 are set, and a 24-hour dynamic update cycle is established to ensure the timeliness of the graph.
[0076] A multi-dimensional conflict detection engine enables accurate identification of compliance risks. The semantic contradiction detector, based on logical rule templates and fuzzy matching technology, sets a detection threshold of 0.7 to identify direct conflicts between clauses. The timeliness conflict analyzer analyzes substitution chain relationships and the validity of transitional clauses by constructing a policy time relationship graph. The authority level verifier establishes a multi-level policy authority system to achieve cross-level conflict arbitration. Finally, by weighted integration of the detection results from each dimension, a quantified conflict feature value is output to support subsequent decision-making. The system deployment utilizes NVIDIA V100 GPU acceleration and is implemented based on the PyTorch framework, supporting batch parallel processing and incremental updates, demonstrating excellent performance in practical applications.
[0077] In S3, a Gaussian mixture model is established to analyze threshold distribution characteristics, a dynamic sliding window algorithm is used to detect outlier data points, a time series prediction model is developed to warn of threshold drift trends, and compliance deviation feature values for pesticide combination characteristics, regional differences, and time dimensions of integrated crops are generated, specifically including:
[0078] The threshold feature library construction adopts a hierarchical classification architecture. First, a three-level classification system (major category - minor category - specific variety) is established according to the "Chinese Pesticide Classification Standard" and the "List of Major Crops," creating an independent feature profile for each crop-pesticide combination. The system accesses nearly 10 years of pesticide use monitoring data from the Ministry of Agriculture and Rural Affairs (over 5 million records). A Gaussian Mixture Model (GMM) is used to fit the threshold distribution, setting 3-5 components to represent typical use intervals. Anomalies (such as exceeding Q3+1.5IQR) are automatically labeled using a semi-supervised learning algorithm, and feature vectors containing 16 statistical measures including mean, standard deviation, and skewness are established. The feature library employs a dual update mechanism: daily incremental updates (processing new data) and triggered major updates (when a significant change in distribution pattern is detected).
[0079] The dynamic anomaly detection system implements an intelligent monitoring strategy adjustment, which specifically includes: a growth cycle perception module that subdivides the crop growth period into stages such as the germination stage (1 - 15 days), the growth stage (16 - 60 days), and the maturity stage (61 - 90 days), and sets detection windows of 7 days, 15 days, and 30 days respectively; a multi-dimensional outlier detection engine that simultaneously considers three dimensions: numerical deviation (Z-score), temporal continuity (DTW distance), and spatial correlation (Moran's I index), and uses the isolation forest algorithm to calculate the anomaly probability; a human-machine collaborative verification process that, when a suspected anomaly is detected, the system automatically associates similar cases in the expert knowledge base (based on Faiss vector retrieval), generates a review report containing historical disposal plans, and after confirmation by an agronomist, feedbacks it to the learning loop.
[0080] The threshold trend prediction system adopts a hybrid architecture of "spatiotemporal attention + LSTM". The technical implementation includes: a data preprocessing layer that performs seasonal decomposition (STL), regional normalization (Z-score normalization), and policy impact annotation on historical threshold data; a feature engineering module that extracts 28 features including seasonal indices (sin / cos encoding), climate features (accumulated temperature, precipitation), and policy change flags; a core prediction model composed of a spatiotemporal attention module (capturing regional associations) and a bidirectional LSTM module (modeling temporal dependencies), and outputs a threshold interval prediction for the next 30 days; a hierarchical early warning mechanism that sets four levels of early warning: blue (predicted value < Q1), yellow (Q1 ≤ predicted value < Q3), orange (Q3 ≤ predicted value < Q3 + 1.5IQR), red (≥ Q3 + 1.5IQR), and triggers different response processes: blue only records logs, yellow initiates enhanced monitoring, orange triggers on-site verification, and red immediately suspends use and initiates a traceability investigation. After the system is deployed, the average discovery time of pesticide use violation incidents is successfully shortened from 17.6 days to 2.3 days, and the early warning accuracy rate reaches 89.7%.
[0081] In S4, the policy conflict feature value and the compliance deviation feature value are fused into a comprehensive risk feature vector, which is input into a multi-modal policy verification model based on a deep residual network. A meta-learning method is used for risk assessment in combination with historical compliance data, and risk traceability analysis is provided through an interpretable AI component, specifically including:
[0082] The multi-modal policy verification model constructed by the present invention adopts an innovative meta-learning architecture, and the specific implementation methods of its core components are as follows:
[0083] The implementation of the feature fusion layer adopts a three-channel attention mechanism:
[0084] The text feature channel uses a pre-trained language model to extract deep semantic representations of policy clauses and automatically calculates weights through semantic relevance analysis, emphasizing the contribution of key constraints. The numerical feature channel is equipped with a credibility evaluator that analyzes the authority of data sources, the standardization of collection methods, and historical accuracy, assigning differentiated weights to values with different levels of reliability. The spatiotemporal feature channel incorporates a timeliness analysis module, considering policy release time, geographical applicability, and seasonal factors to dynamically adjust feature importance. The outputs of the three channels are adaptively fused through a learnable gating mechanism to ensure optimal combination of features from each modality.
[0085] The main architecture of deep residual networks has been specifically optimized:
[0086] The network employs a 50-layer residual structure, with each residual block augmented with agricultural domain knowledge. During initialization, the convolutional kernel parameters are not randomly generated but rather based on a pre-trained model derived from an agricultural policy knowledge graph. Dense connections across layers are incorporated to ensure effective gradient propagation, and adaptive depth supervision is implemented, with auxiliary classifiers added to each intermediate layer. To prevent overfitting, a dynamic stochastic depth technique is used, randomly skipping some residual blocks during training to improve the model's generalization ability.
[0087] The complete workflow of the meta-learning training framework includes:
[0088] First, a structured compliance case knowledge base is constructed, with each case labeled with key features such as policy type, applicable region, and release date, as well as the final decision and basis. The scenario encoder employs a graph neural network architecture, constructing case features as a heterogeneous graph and extracting meta-feature representations of cases through neighbor aggregation and message passing. The parameter prediction network is designed with a hierarchical structure, first generating the architectural parameters of each network module based on the meta-features, and then predicting the specific weight values. The online update system uses a double-buffering mechanism to maintain the stability of the online model while continuously learning new cases in the background, and synchronously updating at fixed intervals.
[0089] Interpretable systems provide comprehensive decision analysis:
[0090] The feature importance analysis module employs a perturbation test method, systematically masking different features to observe and predict changes, quantifying the contribution of each input element. The decision path visualization tool maps the neural network's reasoning process to the correlation paths between policy provisions, generating easily understandable chains of evidence. The comparative analysis function displays the most similar compliance and non-compliance cases, highlighting key differentiating factors. All interpretation results automatically generate structured reports, supporting interactive queries that expand hierarchically.
[0091] The model employs a microservice architecture during deployment, with feature fusion, deep networks, meta-learning, and interpretability components encapsulated as independent services that communicate via message queues. The system undergoes daily incremental training and weekly global model updates to ensure it remains adaptable to the latest policy changes. Practical applications demonstrate that this architecture significantly improves model interpretability and cross-domain adaptability while maintaining high performance.
[0092] In S5, an adaptive optimization strategy is implemented based on the risk assessment results. This includes enhancing the verification strength of high-risk nodes, optimizing negative sampling constraints, and dynamically scheduling computing resources. Through iterative optimization via a closed-loop feedback mechanism, a policy-compliant agricultural knowledge graph supporting compliance monitoring and risk early warning is ultimately generated. Specifically, this includes:
[0093] A tiered resource allocation mechanism establishes a refined risk response system.
[0094] The system sets dynamic risk thresholds (typical values: high risk ≥ 0.8, medium risk 0.5-0.8, low risk < 0.5), and assesses the risk status of each node through a real-time monitoring module. For high-risk nodes, an enhanced verification process is initiated: first, a deep learning-based machine verification engine performs a full analysis; then, the results are pushed to a domain expert workbench for manual review; finally, cross-departmental system comparison confirms consistency. Medium-risk nodes employ a hybrid model of "80% machine verification + 20% random sampling review" to ensure a balance between efficiency and quality. Low-risk nodes are included in a periodic inspection plan, with automated scripts scanning for abnormal indicators daily. All verification results are fed back to the risk scoring model, forming a dynamic adjustment closed loop.
[0095] Intelligent negative sampling system enables adaptive sample management:
[0096] The system maintains a dynamic sampling strategy table, automatically adjusting sampling parameters based on real-time risk assessment results. In high-risk areas (such as nodes related to pesticide bans), the sampling density is increased to three times the baseline value to ensure full coverage of potential problems. Protective constraints are set for policy-sensitive relationships (such as cross-regional application clauses) to avoid generating invalid negative samples. The system integrates an adversarial example generator to construct challenging training samples by adding semantic perturbations and numerical noise. The sampling effectiveness evaluation module continuously monitors the model's performance on the validation set and automatically adjusts the sampling strategy parameters.
[0097] Elastic resource management adopts a cloud-native architecture:
[0098] A hybrid cloud resource pool is built, integrating local GPU servers and public cloud elastic computing resources. The task scheduler automatically categorizes tasks based on their content: critical tasks (such as high-risk node verification) are prioritized for GPU resource allocation; routine tasks are processed using CPU clusters; and batch jobs are scheduled to low-cost computing resources. A parallel processing framework enables fine-grained task partitioning and supports multi-level pipeline execution. A resource monitoring dashboard displays the load status of each node in real time, automatically triggering horizontal scaling when utilization exceeds a threshold.
[0099] Self-optimizing closed-loop system construction and continuous evolution capability:
[0100] The data acquisition agent captures policy execution metrics in real time, including verification accuracy, processing latency, and resource consumption. The analysis engine employs association rule mining and causal inference techniques to identify potential patterns between policy parameters and their effects. The parameter optimizer automatically adjusts key parameters such as thresholds and sampling rates based on the analysis results, verifying the effectiveness of each change through A / B testing. Validated adjustments are transformed into training data, driving iterative updates to the system model. The system generates monthly optimization reports, showcasing trends in key metrics and improvement methods.
[0101] The working principle of this invention is as follows: This invention achieves precise control over policy compliance through a five-stage collaborative processing flow. First, a distributed data acquisition system is constructed, employing adaptive crawling technology to acquire structured policy texts, numerical threshold standards, and logical constraint rules in real time, and ensuring data quality through multi-level verification. Second, a hierarchical semantic parsing architecture is designed, utilizing bidirectional LSTM networks and graph attention networks to extract deep features from policy texts, constructing a dynamically updated semantic association graph, and generating quantified conflict features by combining timeliness factors. Next, a threshold analysis and prediction system is established, identifying abnormal patterns through a Gaussian mixture model and using a spatiotemporally aware LSTM model to warn of compliance risks. Subsequently, a multimodal fusion verification model is developed, integrating deep residual networks and a meta-learning framework for comprehensive risk assessment, and equipped with interpretable components to achieve decision tracing. Finally, an intelligent optimization strategy is implemented, dynamically adjusting verification intensity, sampling strategies, and resource allocation based on risk levels to form a continuously evolving knowledge graph construction closed loop. This system innovatively integrates deep learning and domain knowledge, achieving a conflict detection accuracy rate of 92.3% in agricultural policy analysis scenarios and improving the timeliness of violation detection by 85%, significantly enhancing the compliance and practicality of agricultural knowledge graphs.
[0102] The above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more sets of available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium. A semiconductor medium can be a solid-state drive.
[0103] It should be understood that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. A and B can be singular or plural. Additionally, the character " / " in this article generally indicates an "or" relationship between the preceding and following related objects, but it can also represent an "and / or" relationship. Please refer to the context for a more accurate understanding.
[0104] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0105] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.
Claims
1. A method for constructing agricultural knowledge graphs based on large models, characterized in that, Includes the following steps: S1: Construct a distributed data acquisition system to acquire structured policy texts, numerical threshold standards, and logical constraint rules in real time. Dynamically capture updates of multi-source agricultural policy data through adaptive crawling technology, and adopt a data quality verification mechanism to ensure the integrity and consistency of input data. S2: A bidirectional LSTM network based on an attention mechanism is used to extract deep semantic features of policy texts, construct a semantic association graph of policy clauses, use a graph neural network to calculate the semantic distance between nodes, and dynamically adjust the conflict judgment threshold by combining the timeliness weight factor to generate quantitative policy conflict feature values. S3: Establish a Gaussian mixture model to analyze the threshold distribution characteristics, use a dynamic sliding window algorithm to detect abnormal data points, develop a time series prediction model to warn of threshold drift trends, and generate compliance deviation characteristic values of pesticide combination characteristics, regional differences, and time dimensions for integrated crops. S4: The policy conflict feature value and the compliance deviation feature value are integrated into a comprehensive risk feature vector, which is input into a multimodal strategy verification model based on deep residual network. The meta-learning method is used to conduct risk assessment in combination with historical compliance data, and risk source analysis is provided through interpretable AI components. S5: Implement adaptive optimization strategies based on risk assessment results, including enhancing the verification strength of high-risk nodes, optimizing negative sampling constraints, and dynamically scheduling computing resources. Through iterative optimization via a closed-loop feedback mechanism, a policy-compliant agricultural knowledge graph that supports compliance monitoring and risk early warning is ultimately generated.
2. The method for constructing an agricultural knowledge graph based on a large model according to claim 1, characterized in that, The process for obtaining the policy conflict characteristic values is as follows: A hierarchical semantic parsing framework is used to process policy texts. First, basic semantic features are extracted by word-level bidirectional LSTM. Then, paragraph-level attention mechanism is used to focus on key clauses. Finally, document-level graph neural network is used to model the global relationship between policy clauses. Construct a dynamically updated policy semantic association graph, where nodes represent policy clauses, and edge weights are jointly determined by semantic similarity and timeliness similarity. Semantic similarity is calculated based on deep semantic matching of clause content, while timeliness similarity considers the time interval between policy releases and revision history. A multi-dimensional conflict detection algorithm is designed to analyze the semantic contradictions, timeliness conflicts, and differences in authority levels of policy provisions, and output a weighted conflict score matrix.
3. The method for constructing an agricultural knowledge graph based on a large model according to claim 2, characterized in that, The hierarchical semantic parsing specifically includes: In the word-level feature extraction stage, a bidirectional LSTM network is used to capture the contextual semantic information of the policy text; The paragraph-level processing stage uses a multi-head attention mechanism to identify key constraints in the clauses; In the document-level integration phase, semantic relationships across clauses are established using graph attention networks; The final policy text representation combines local details with global structure.
4. The method for constructing an agricultural knowledge graph based on a large model according to claim 2, characterized in that, The multi-dimensional conflict detection specifically includes: Semantic contradiction detection: Analyze whether there are direct conflicts in the logical constraints between clauses; Timeliness conflict assessment: Determine the substitution relationship and transitional provisions between the old and new policies; Authority level verification: Different conflict determination standards are set for policies at different levels; Comprehensive score generation: The final conflict feature value is output by weighted integration of the detection results from various dimensions.
5. The method for constructing an agricultural knowledge graph based on a large model according to claim 1, characterized in that, The process for obtaining the compliance deviation feature value is as follows: Construct a threshold feature library for crop and pesticide combinations, establish independent threshold distribution models for different crop types and pesticide varieties, and analyze the standard distribution patterns of usage in historical data through machine learning algorithms to identify typical threshold features of various combinations; A dynamic anomaly detection mechanism is implemented, which automatically adjusts the detection window size according to the crop growth cycle stage, uses multi-dimensional outlier identification technology to find abnormal data points, and combines it with an expert knowledge base for manual review and confirmation. Develop a threshold trend prediction system that integrates time series analysis and deep learning technologies to establish a prediction model that takes into account factors such as seasonal changes, regional differences, and policy adjustments, and provide early warnings of potential compliance risks.
6. The method for constructing an agricultural knowledge graph based on a large model according to claim 5, characterized in that, The threshold feature library specifically includes: Establish a classification system based on crop type and pesticide type; Collect historical usage data to build a distribution model; Mark typical threshold ranges and abnormal features; Configure an automatic model update mechanism.
7. The method for constructing an agricultural knowledge graph based on a large model according to claim 5, characterized in that, The potential compliance risks mentioned in the warning specifically include: The detection period is divided according to the crop growth cycle. Multi-scale sliding window technology is used, combined with rule engine and machine learning to identify anomalies and establish an anomaly data hierarchical processing flow. Historical threshold change data are collected, seasonal and regional characteristics are analyzed, and a deep learning prediction model is constructed. Based on the output of the deep learning prediction model, a multi-level early warning mechanism is set up. The first level is to quickly screen for threshold overruns based on the statistical model. The second level is to identify abnormal patterns by combining spatiotemporal characteristics. The third level introduces an expert knowledge base for in-depth verification. The fourth level conducts cross-regional risk correlation analysis.
8. The method for constructing an agricultural knowledge graph based on a large model according to claim 1, characterized in that, The construction of the multimodal strategy verification model includes the following steps: The design incorporates a feature fusion layer that dynamically fuses policy conflict feature values and compliance deviation feature values using an attention-weighted mechanism. Assign semantic weights to text features; Assign confidence weights to numerical features; Assign timeliness weights to spatiotemporal features; We construct the main architecture of a deep residual network, adopt cross-layer connection to avoid gradient vanishing, and embed a domain knowledge-guided convolutional kernel initialization strategy in each residual block. Develop a meta-learning training framework to extract meta-features from historical compliance cases and build a context-aware model parameter rapid adaptation mechanism; Integrating interpretable AI components enables traceability of the risk assessment process through feature importance analysis and decision path visualization.
9. The method for constructing an agricultural knowledge graph based on a large model according to claim 1, characterized in that, The meta-learning training framework specifically includes: Build a multi-granularity compliance case library, marking the key features and outcomes of each case; Design a scenario encoder to extract meta-feature representations of cases; Develop a parameter prediction network to dynamically generate model parameters based on the meta-features of new cases; Establish an online update mechanism to continuously optimize the meta-knowledge representation.
10. The method for constructing an agricultural knowledge graph based on a large model according to claim 1, characterized in that, The implementation of the adaptive optimization strategy specifically includes: Establish a hierarchical resource scheduling mechanism to dynamically allocate verification resources based on node risk scores, specifically including: If a node's risk score is greater than or equal to the first preset threshold, it is recorded as a high-risk node, and a three-level verification process is initiated, including machine verification, expert review, and cross-system comparison. If a node's risk score is less than the first preset threshold but greater than or equal to the second preset threshold, it is designated as a medium-risk node, and a combination of machine verification and sampling review is used. If the node risk score is less than the second preset threshold, it is recorded as a low-risk node and a regular inspection mechanism is implemented. Develop an intelligent negative sampling optimization system that automatically adjusts the sampling strategy based on risk characteristics, increases sampling density in high-risk areas, sets special sampling constraints for policy-sensitive relationships, introduces adversarial sample generation technology to improve model robustness, and establishes a continuous evaluation mechanism for sampling effects. Implement a flexible resource allocation scheme, build a dynamically scalable computing resource pool, implement priority scheduling according to the criticality of tasks, develop a parallel processing framework that supports load balancing, and monitor resource utilization efficiency indicators in real time. A self-optimizing closed-loop system is constructed to continuously collect data on the effectiveness of strategy execution, deeply analyze the correlation between strategy parameters and optimization results, automatically adjust the configuration parameters of the optimization strategy, and generate new training datasets through effect verification to drive the continuous evolution of the system.
Citation Information
Cited By
Large model auxiliary decision-making method and system applied to natural resource informatization management
CN121615755A
Equal-security compliance knowledge base construction method based on multi-dimensional semantic association
CN122133776A