Method and system for constructing multi-modal knowledge graph in agricultural field
By constructing an agricultural ontology and evidence templates for semantic matching and optimization of multimodal knowledge graphs, the problems of credibility and dynamic updating of knowledge graphs in existing technologies are solved. This enables the unified semantic expression of multimodal data and the dynamic evolution of knowledge graphs, thereby improving the accuracy and timeliness of agricultural knowledge.
Patent Information
- Application Number
- CN202511672730.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-14
- Publication Date
- 2026-02-10
AI Technical Summary
Existing methods for constructing multimodal knowledge graphs in agriculture lack systematic modeling and closed-loop optimization mechanisms for cross-modal alignment credibility, ontology constraints, evidence tracing, causal verification, and dynamic updates. This leads to knowledge becoming distorted over time, making it difficult to achieve verifiability and intelligent usability.
We collect multi-source agricultural data to construct an agricultural ontology, perform semantic matching and optimization through agricultural semantic rules and evidence templates, generate structured knowledge units, establish an event causal relationship network, and perform incremental correction and self-correction when new data arrives or conflicts occur. We then combine evidence and causal information to perform reasoning and recommendations.
It achieves unified semantic expression of multimodal data such as text and images, improves the reliability of knowledge extraction results, and enhances the accuracy, interpretability and timeliness of agricultural knowledge organization through dynamic evolution and adaptive updates, providing technical support for intelligent agricultural decision-making.
Smart Images

Figure CN121502012A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and in particular to a method and system for constructing a multimodal knowledge graph in the agricultural field. Background Technology
[0002] With the development of smart agriculture, agricultural knowledge graphs have become an important means to achieve the structuring of agricultural information and the intelligent management of knowledge. Existing technologies mainly describe entities and relationships in the agricultural field using triples, and combine them with methods such as feature extraction and semantic matching to achieve structured management of knowledge such as crop characteristics, agricultural tool types, and agricultural activity stages. Some studies have introduced multimodal fusion technology combining text and images, using attention mechanisms to associate information from different modalities, providing a new direction for the automatic construction of agricultural knowledge. However, these methods mostly remain at the descriptive level, primarily achieving information fusion and display, and are still insufficient to fully express the dynamic characteristics and temporal changes of agricultural production behaviors.
[0003] In recent years, research trends in agricultural intelligence have gradually shifted from data availability to data reliability and model credibility. Knowledge graph construction is moving towards interpretability, verifiability, and traceability. Agricultural production activities involve complex temporal sequences and causal logic, with frequent changes in external factors and a high degree of data diversity. Therefore, academia and industry are increasingly emphasizing research on multimodal knowledge graphs with scalability and dynamic update capabilities, hoping to maintain semantic consistency and structural stability of knowledge graphs during data updates and model iterations through incremental learning, cross-modal alignment, and self-supervised optimization. Simultaneously, the semantic representation of agricultural knowledge is expanding from static entity relationships to dynamic modeling of events and causal chains, thereby better supporting field management, agricultural reasoning, and precise decision-making.
[0004] However, existing technologies still have several shortcomings. First, the credibility of multimodal alignment has not been effectively quantified, and the model output lacks a confidence assessment mechanism, making it difficult to measure the reliability of the fusion results. Second, existing methods lack guidance from agricultural ontology and constraint rules, easily leading to problems such as role confusion, stage errors, and improper event connections during semantic extraction. Third, the credibility and time freshness of data sources are not considered during knowledge fusion, making it difficult to achieve conflict resolution and evidence tracing. Although existing graphing methods introduce temporal consistency constraints, they remain at a simple guarantee level of sequence, failing to achieve systematic verification and counterfactual analysis of causal logic. In addition, knowledge graphs lack incremental updates and closed-loop feedback mechanisms, making it impossible to dynamically repair and self-correct when new data arrives, resulting in knowledge gradually becoming distorted over time. Finally, inference results largely rely on path probability calculations, failing to comprehensively consider uncertainty propagation and source reputation, thus exhibiting significant limitations in credibility assessment and recommendation explanation. These problems collectively constrain the verifiability and intelligent usability of multimodal knowledge graphs in the agricultural field. Summary of the Invention
[0005] To overcome the shortcomings of existing technologies, the purpose of this invention is to provide a method and system for constructing multimodal knowledge graphs in the agricultural field. This invention solves the problem that existing methods for constructing multimodal knowledge graphs in agriculture lack system modeling and closed-loop optimization mechanisms for cross-modal alignment credibility, ontology constraints, evidence tracing, causal verification, and dynamic updates.
[0006] To achieve the above objectives, the present invention provides the following solution: a method for constructing a multimodal knowledge graph in the agricultural field, comprising: collecting multi-source agricultural data and analyzing the entity types, semantic relationships, and agricultural activity sequences of the multi-source agricultural data to construct an agricultural ontology; constructing agricultural semantic rules and evidence templates based on the agricultural ontology; performing semantic matching on the text features and image features of the multi-source agricultural data according to the agricultural ontology and evidence templates to obtain preliminary fusion features; optimizing the preliminary fusion features under ontology constraints to obtain a fusion semantic representation and corresponding confidence results; identifying agricultural entities, relationships, and events based on the fusion semantics and corresponding confidence results, and performing credibility determination and conflict resolution based on evidence information to generate structured knowledge units; constructing an event causal association network based on the structured knowledge units; generating an agricultural multimodal knowledge graph based on the event network, and performing incremental correction and self-correction when new data arrives or conflicts occur to obtain an updated knowledge graph; and performing reasoning and recommendation based on the updated graph, combined with evidence and causal information, and outputting agricultural auxiliary decision-making results with confidence.
[0007] The present invention discloses the following technical effects:
[0008] This invention provides a method and system for constructing a multimodal knowledge graph in the agricultural field. The method includes: collecting multi-source agricultural data and analyzing the entity types, semantic relationships, and agricultural activity sequences of the multi-source agricultural data to construct an agricultural ontology; constructing agricultural semantic rules and evidence templates based on the agricultural ontology; performing semantic matching on the text features and image features of the multi-source agricultural data according to the agricultural ontology and evidence templates to obtain preliminary fusion features; optimizing the preliminary fusion features under ontology constraints to obtain fused semantic representations and corresponding confidence results; identifying agricultural entities, relationships, and events based on the fused semantics and corresponding confidence results, and performing credibility judgment and conflict resolution based on evidence information to generate structured knowledge units; constructing an event causal association network based on the structured knowledge units; generating an agricultural multimodal knowledge graph based on the event network, and performing incremental correction and self-correction when new data arrives or conflicts occur to obtain an updated knowledge graph; and performing reasoning and recommendation based on the updated graph, combined with evidence and causal information, and outputting agricultural auxiliary decision-making results with confidence. This invention achieves unified semantic expression of multimodal data such as text and images through ontology-driven semantic modeling and fusion feature optimization; it improves the reliability of knowledge extraction results by utilizing confidence judgment and conflict resolution; and it realizes dynamic evolution and adaptive updating of the knowledge graph through the establishment of an event causal association network and an incremental self-correction mechanism. This effectively solves the problems of knowledge update lag, missing causal relationships, and low credibility of reasoning results in traditional methods, and can significantly improve the accuracy, interpretability, and timeliness of agricultural knowledge organization, providing technical support for intelligent agricultural decision-making and knowledge services. Attached Figure Description
[0009] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0010] Figure 1 A flowchart illustrating a method for constructing a multimodal knowledge graph in the agricultural field, provided as an embodiment of the present invention. Detailed Implementation
[0011] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0012] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0013] like Figure 1 As shown, this invention provides a method for constructing a multimodal knowledge graph in the agricultural field, including:
[0014] Step 100: Collect multi-source agricultural data and analyze the entity types, semantic relationships, and agricultural activity sequences of the multi-source agricultural data to construct an agricultural ontology;
[0015] Step 200: Construct agricultural semantic rules and evidence templates based on the agricultural ontology;
[0016] Step 300: Based on the agricultural ontology and evidence template, perform semantic matching on the text features and image features of the multi-source agricultural data to obtain preliminary fusion features;
[0017] Step 400: Optimize the preliminary fusion features under ontology constraints to obtain the fusion semantic representation and the corresponding confidence results;
[0018] Step 500: Identify agricultural entities, relationships, and events based on the fused semantics and corresponding confidence results, and determine credibility and resolve conflicts based on evidence information to generate structured knowledge units;
[0019] Step 600: Construct an event causal relationship network based on structured knowledge units;
[0020] Step 700: Generate an agricultural multimodal knowledge graph based on the event network, and perform incremental correction and self-correction when new data arrives or conflicts occur to obtain an updated knowledge graph;
[0021] Step 800: Based on the updated graph, reason and recommend by combining evidence and causal information, and output agricultural auxiliary decision-making results with confidence.
[0022] Furthermore, the specific implementation process of step 100 is as follows:
[0023] This embodiment first addresses the challenges of complex data sources and diverse formats in the agricultural field by employing multimodal data unification and cleaning techniques to preprocess multi-source agricultural data from agricultural monitoring platforms, operational records, and remote sensing imagery. Through automatic format conversion, noise filtering, and data alignment, it ensures semantic interoperability between text and image data, providing high-quality input for subsequent semantic extraction. The processed multi-source agricultural data encompasses agricultural literature, agricultural machinery operation images, crop growth monitoring images, and expert descriptions, significantly improving data consistency and usability.
[0024] In this embodiment, after data cleaning, a combination of a pre-trained language model and an agricultural domain dictionary is used to perform semantic segmentation and named entity recognition on the corpus. This extracts core entities related to agricultural production, including crop types, agricultural machinery, environmental factors, and agricultural activities. These entities are then hierarchically labeled based on semantic hierarchy features to form an agricultural entity set. During this process, precise classification is achieved through the strength of contextual dependencies and concept distance between entities to ensure the semantic correctness of subsequent relation extraction.
[0025] This embodiment, based on the spatiotemporal characteristics and causal logic of agricultural production, employs a semantic parsing algorithm based on dependency syntax and a time series analysis algorithm to extract temporal, causal, and dependency relationships from a set of agricultural entities, forming a semantic relation set. Combining the stage characteristics of agricultural operations, temporal analysis is performed on the semantic relation set to generate a temporal structure of agricultural activities. Based on the triple constraints of entities, relations, and temporal structure, a semantically complete agricultural ontology is constructed, achieving a systematic modeling of the semantic framework of agricultural knowledge.
[0026] Furthermore, the specific implementation process of step 200 is as follows:
[0027] This embodiment first extracts entity types, attribute relationships, and temporal and causal logic constraints from the constructed agricultural ontology, organizing them into a machine-readable semantic constraint set. This set, through ontology hierarchical relationships and constraint dependencies, clarifies the structural hierarchy and interaction limitations of each entity type in the agricultural production process, providing logical boundary conditions for subsequent rule generation. During the extraction process, this embodiment uses agricultural crops, agricultural machinery, environmental factors, and agricultural activities as core concepts, analyzing object attributes and data attributes in the ontology to obtain constraint patterns that reflect the semantics of real-world operations.
[0028] This embodiment automatically derives agricultural semantic rules based on the aforementioned set of semantic constraints using a rule generation algorithm. The algorithm traverses the semantic graph structure of entities and attributes, identifies behavioral dependency chains and conditional trigger chains, and outputs rule expressions in the form of logical predicates, forming a semantic logic set that describes "several conditions → events → results." During the generation process, the algorithm automatically integrates the process characteristics of each stage of agricultural production, ensuring that the generated rules not only contain static constraints but also reflect the temporal and causal logic of agricultural activities, guaranteeing the consistency and scalability of the rule system in reasoning and verification.
[0029] Based on the data relationships and feature types involved in the aforementioned semantic rules, this embodiment defines evidence fields to identify data sources, collection times, collection methods, and confidence levels, and constructs an agricultural evidence template accordingly. This template serves as a unified standard for data credibility and traceability, mapping the source characteristics of multi-source data to the evidence information structure through field correspondences. This achieves the association and binding of semantic rules and data evidence, thereby supporting the quantitative determination of evidence reliability during subsequent knowledge extraction, verification, and conflict resolution processes.
[0030] Specifically, in this embodiment, when generating agricultural semantic rules, a semantic graph traversal is first performed on the entity nodes and their attribute relationships in the semantic constraint set. Using a depth-first parsing approach, the constraint directions and dependency paths between entities are identified sequentially. Node pairs that satisfy semantic consistency and ontology constraints are marked, forming candidate behavior chains. These candidate behavior chains are centered on the interaction methods of entities, recording participating elements, constraints, and triggering variables, providing a foundation for subsequent logical deduction. This operation ensures that the entity relationship extraction during rule generation has structural integrity and contextual traceability.
[0031] This embodiment, based on the construction of behavioral dependency chains, introduces a temporal analysis mechanism according to the stage characteristics of the agricultural production process, mapping agricultural activities with stage characteristics onto a timeline. When traversing each behavioral chain, the algorithm compares the temporal order of activities, participating objects, and environmental conditions to automatically identify the sequential logic and necessary conditions for triggering behaviors, determining the action sequence with causal relationships, and using this sequence as the main body of the conditional trigger chain. During the formation of the conditional trigger chain, weak trigger relationships without temporal correlation are automatically excluded, ensuring that the generated logical chain strictly conforms to the production logic of agricultural operations.
[0032] This embodiment identifies the behavioral dependency chain and the condition trigger chain, and then transforms them into logical predicate expressions through a unified logic combination module. Each rule adopts the logical structure of "if the set of conditions is true, then an event occurs and a result is caused," consisting of antecedent conditions, triggering events, and consequent results, outputting a standardized set of rule expressions. During the generation process, the algorithm automatically controls the rule complexity and applicable domain based on ontology constraints, ensuring that the formed agricultural semantic rules have good reasoning operability and semantic consistency, thereby achieving high-precision logical judgment in subsequent knowledge reasoning and consistency verification stages.
[0033] Furthermore, the specific implementation process of step 300 is as follows:
[0034] This embodiment first targets textual agricultural data, utilizing a semantic embedding-based feature extraction model to extract semantic information. By segmenting agricultural literature, operational records, and expert descriptions into sentences and performing word form normalization, keywords and entity labels related to agricultural entities are identified, and corresponding word vector representations are generated. Simultaneously, a context window modeling method is used to capture semantic environment information, mapping the entity association strength within the same semantic segment to contextual semantic feature vectors. After feature regularization, a text feature set containing word vectors, entity labels, and contextual semantics is formed, ensuring that the text representation is distinguishable and relevant in a multi-dimensional semantic space.
[0035] This embodiment focuses on agricultural image data, applying target detection and feature embedding extraction algorithms to perform region segmentation and target identification on crop growth images, agricultural machinery operation images, and remote sensing monitoring images. Visual feature vectors and scene labels are extracted from the identified target regions. Visual features include attributes related to agricultural objects, such as color distribution, shape contours, and texture information. Preliminary semantic alignment is achieved by matching image regions with agricultural ontological entity types, giving the image features clear semantic directionality. After dimensionality normalization and feature filtering, an image feature set corresponding to the text features is obtained.
[0036] This embodiment performs semantic mapping and matching of text and image features based on entity categories and semantic relationships in the agricultural ontology. First, a semantic space alignment algorithm is used to calculate the similarity between text word vectors and image visual vectors in the same semantic dimension. This similarity is then weighted by combining confidence weights from the evidence template to obtain a cross-modal matching degree matrix. Subsequently, features are adaptively aligned based on the matching degree matrix, fusing semantically consistent text and image features to form a weighted fusion preliminary fusion feature. This achieves a unified semantic expression of text and images, providing high-quality input for subsequent fusion optimization and knowledge extraction.
[0037] Furthermore, the specific implementation process of step 400 is as follows:
[0038] Before optimizing the initial fused features, this embodiment first constructs an ontology constraint set for feature optimization based on the hierarchical relationships and attribute constraints of the agricultural ontology. This embodiment extracts consistency conditions and hierarchical constraint rules reflecting the logic of agricultural knowledge by analyzing entity types, attribute dependencies, and semantic relationships in the agricultural ontology, and transforms them into a computable embedding matrix form. This ontology constraint set is used to limit the distribution boundaries of the fused features in the semantic space, ensuring that features of different modalities have consistent semantic orientation and logical constraint relationships within the same semantic context during the optimization process, providing a standardized basis for subsequent feature weight adjustment and consistency correction.
[0039] This embodiment, after obtaining the ontology constraint set, performs semantic consistency adjustment and feature weight reallocation on the preliminary fused features. By calculating the semantic deviation between each fused feature and the ontology constraints, its conformity with the agricultural semantic hierarchy is evaluated, and the weight ratio of the feature components is adaptively adjusted based on the deviation. For fused features with synonymous or near-synonymous semantic conflicts, this embodiment introduces a semantic offset compensation mechanism to adjust their vector direction to ensure feature synergy in the multimodal space. Through this process, the adjusted fused features possess both semantic consistency and structural stability, laying the foundation for further semantic optimization.
[0040] This embodiment then performs multimodal semantic optimization and confidence calculation. Through a cross-modal semantic preservation algorithm, global relevance is maintained between the text and image semantic spaces, reducing feature distortion caused by modality shifts to generate a semantically stable fused representation. Next, based on the data source, time attribute, and credibility weight information in the evidence template, this embodiment evaluates the matching strength and ontology consistency of the fused semantic representation, calculating the confidence value for each fused semantic unit. The confidence results are normalized to ensure they fall within a uniform evaluation range, ultimately yielding the fused semantic representation and corresponding confidence results, providing a reliable basis for subsequent entity recognition and knowledge verification.
[0041] Furthermore, the specific implementation process of step 500 is as follows:
[0042] Before identifying agricultural entities, relationships, and events, this embodiment first performs structured semantic parsing on the fused semantic representation. A semantic structure analysis algorithm is used to hierarchically deconstruct the fused semantic representation, extracting the dependency structure, hierarchical logic, and semantic orientation between semantic units to construct a tree-like semantic dependency model. This embodiment combines multi-level syntactic constraints in agricultural semantic rules to cluster and label semantic nodes according to their semantic function categories, distinguishing between entity, attribute, and action terms, and establishing a cross-modal semantic matching index. After processing, an agricultural semantic representation model with clear semantic hierarchical relationships and a parsable structure is formed, providing a high-precision input foundation for subsequent knowledge element identification.
[0043] After obtaining the agricultural semantic representation model, this embodiment automatically establishes an entity identifier and attribute set that conforms to the agricultural semantic system based on the entity categories and attribute ranges defined in the agricultural ontology. This process uses semantic pattern matching to determine the entity affiliation in text and image information, identifying its category, attributes, and state characteristics, and outputting standardized agricultural entity nodes. Simultaneously, based on agricultural semantic rules, semantic dependency edges are identified between the identified entities, extracting a set of relationships and events reflecting production logic and operational sequence. At this stage, this embodiment uses temporal constraints and causal logic judgments to ensure that the triggering conditions and causal results of events conform to agricultural activity procedures, achieving interpretable agricultural event identification.
[0044] After extracting entities and relationships, this embodiment performs credibility matching and conflict resolution on entity, relationship, and event sets based on the source, timestamp, and confidence weight information in the evidence template. First, it calculates the confidence difference and semantic consistency index between the evidence corresponding to each knowledge element, thus assigning a credibility level. Then, it filters out knowledge elements with confidence or semantic conflicts, executes a weighted conflict resolution mechanism, and selects the optimal knowledge representation by comparing the timeliness and source credibility of each element. Finally, this embodiment structurally encapsulates the agricultural entities, relationships, and events after conflict resolution and confidence assessment, generating structured knowledge units containing semantic, relational, and event metadata, thus achieving standardized and reasonable expression of agricultural knowledge.
[0045] Specifically, in constructing the agricultural semantic representation model, this embodiment achieves the joint expression of agricultural entities, relationships, event contexts, and evidence confidence by performing nonlinear semantic mapping operations on the fused features. Specifically, this embodiment first inputs the agricultural entity feature vector, relationship representation vector, event context feature vector, and confidence weight factor into the semantic fusion layer. The semantic fusion layer embeds a proportional parameter control module with adaptive learning capabilities to dynamically adjust the influence weight of each semantic component in the overall semantic representation. Through feature standardization and nonlinear normalization mapping, features from different semantic sources achieve a consistent scale and semantic orientation in the same embedding space, thereby forming a comprehensive fused representation of agricultural semantic units.
[0046] Agricultural entity features are represented by vectorized representations of entities such as crops, farm tools, environmental elements, and agricultural activities after fusion and optimization, reflecting the semantic attributes of agricultural subjects; relationships are represented by semantic vectors describing the dependency, causal, or temporal logic between different agricultural entities, reflecting the semantic connections between entities; event context features are semantic representations of the event environment and process state extracted from fused semantics, used to reflect the semantic scenarios in which agricultural behavior occurs; confidence weight factors are parameters calculated based on agricultural evidence templates, used to assess the reliability of evidence for corresponding semantic units.
[0047] In setting the adaptive scaling parameters, this embodiment uses an internal gradient update algorithm to dynamically learn each parameter, automatically adjusting the weights of different semantic components in the overall fusion representation according to data distribution and semantic complexity. For example, when agricultural semantic units originate from multimodal cross-information, this embodiment increases the proportional weights of entity and relation features through iterative optimization, thereby strengthening the robustness of semantic reasoning. When semantic conflicts or significant differences in evidence credibility are detected, the weight ratio of the confidence factor is automatically increased to maintain the reliability of the fusion result. Through the above methods, the final generated agricultural semantic fusion representation can accurately reflect the comprehensive semantic structure of agricultural knowledge under multi-source data conditions and possesses good interpretability and scalability.
[0048] More specifically, firstly, the agricultural entity feature embedding is obtained by fusing and optimizing multi-source agricultural data. For example, given the text "a description of winter wheat needing field irrigation during the grain-filling stage" and the corresponding field image sample, this embodiment obtains the feature embedding vector of the entity "winter wheat" through a feature extraction algorithm. Its main dimension is 256, with each dimension representing semantic information such as growth stage, crop type, and environmental conditions. The numerical range of this feature after fusing and optimization is typically between -1.0 and 1.0; for example, a set of feature values for entity embedding could be 0.46, -0.22, 0.78, -0.15, and 0.53.
[0049] Secondly, the semantic representation of the relationships between entities comes from the semantic parsing module's mining of the semantic dependencies between entities. For example, in the case above, "irrigation" as an agricultural activity forms a "crop-water requirement-irrigation behavior" relationship with the entity "winter wheat". After being vectorized by the relation encoder, this relationship generates a relation vector of length 128, with typical eigenvalues distributed between -0.5 and 0.5, such as -0.14, 0.27, -0.31, 0.42, and -0.09.
[0050] Secondly, the event context features are derived from the aggregation and extraction of scene and time information by the multimodal semantic analysis model. Continuing with the example of the "irrigation during the grouting stage" event, its context features include "time: grouting period," "scene: field," "weather: sunny," and "temperature: 25℃," etc. After encoding, a context vector of length 64 is formed, with values generally ranging from 0 to 1, such as 0.63, 0.88, 0.21, 0.47, and 0.75.
[0051] The confidence weight factor comes from the credibility measurement module of the agricultural evidence template. For example, this feature is determined by the reliability of the data source (remote sensing monitoring, expert records, on-site image recognition, etc.). When the data comes from an official agricultural monitoring platform, the confidence level can be set to 0.92; when the source is manual records or experience descriptions, the confidence level can be 0.68; when multiple sources are fused, a weighted average is taken, such as (0.92×0.7 + 0.68×0.3)=0.84, then the confidence weight of this semantic unit is 0.84.
[0052] The proportionality coefficient is an adjustment parameter automatically learned by the model, controlling the contribution ratio of each feature to the final semantic fusion result. Taking a single training result as an example, after multiple rounds of gradient updates, the system obtains a first proportionality coefficient of 0.41, a second proportionality coefficient of 0.36, a third proportionality coefficient of 0.17, and a fourth proportionality coefficient of 0.06. This can be understood as entity semantics playing a dominant role of 41% in the overall semantic fusion, relations 36%, event context 17%, and confidence adjustment 6%.
[0053] The nonlinear semantic normalization mapping function originates from a multi-layer semantic mapping network. It typically employs a nonlinear transformation with a modified linear unit to normalize and compress the embedded vectors, limiting the output to the interval between -1 and 1. For example, after mapping, a portion of the fusion results from the input combined vectors are 0.73, -0.44, 0.59, -0.22, and 0.81, forming the final fusion representation of the agricultural semantic unit, which is used for subsequent agricultural knowledge recognition and reasoning calculations.
[0054] Furthermore, the specific implementation process of step 600 is as follows:
[0055] In constructing the event causal relationship network, this embodiment first uses pre-generated structured knowledge units as input data sources, extracting semantic vectors, occurrence times, and corresponding entity relationship information representing events. Through semantic coupling modeling, the logical relevance and temporal constraints between events are calculated within a unified semantic space, ultimately forming a causal relationship network to describe the logical relationships of mutual influence and triggering of events in agricultural production. The core idea of this network is to comprehensively consider semantic similarity, common causal features, temporal relationships, and confidence differences, ensuring that the causal chain possesses both statistical relevance and logical rationality within agricultural knowledge.
[0056] In this implementation, the semantic feature vector of each event is input into the causal mapping module. The module contains an automatically learnable causal mapping matrix to measure the influence weights between the semantic representations of events. The common-cause similarity between events is derived from the analysis of shared entities and common environmental elements. For example, if the events "irrigation behavior" and "disease reduction" share the same crop and similar climatic conditions, their common-cause similarity is 0.82; if the two events occur in different crops and different environments, the similarity can be reduced to 0.24. The time interval coefficient is obtained from the timestamp difference in the event record data, with values in days. For example, the time difference for events occurring on the same day is 0.1, the time difference for two adjacent days on different days is 0.5, and the difference spanning several weeks can reach 0.9, used to quantify the temporal constraints between events.
[0057] Furthermore, this embodiment introduces two adaptive adjustment parameters and a nonlinear normalization function to balance the importance of semantics and temporal sequence. The first adjustment parameter adjusts the weight of common cause similarity in the overall causal strength; for example, a value of 0.35 indicates that common cause features contribute 35% to the overall association. The second adjustment parameter controls the inhibitory effect of time difference on causal inference; for example, a value of 0.22 indicates that the greater the temporal distance, the more significant the attenuation of the association strength. The nonlinear normalization function limits the output causal strength result to between 0 and 1, so that the model can converge stably and is easy to interpret. For example, in one calculation, the semantic relationship value between the events "spring fertilization" and "increased crop yield" is 0.76, the common cause similarity is 0.82, and the time interval is 0.3. After the above mapping and adjustment, the output causal association strength is 0.68. This result shows that there is a significant positive causal relationship between the two events, which can be incorporated into the agricultural event causal network for inferring the optimal path and decision-making basis of agricultural production activities.
[0058] Furthermore, the specific implementation process of step 700 is as follows:
[0059] In generating an agricultural multimodal knowledge graph, this embodiment first constructs an initial graph structure in a graph database based on an event causal relationship network, with agricultural entities as nodes and relationships and events as edges. This embodiment assigns a globally unique identifier to each agricultural entity through a node identifier generation module and determines the types of nodes and edges according to agricultural semantic rules. For example, "rice," "irrigation event," and "nitrogen fertilizer application" are used as nodes, while semantic relationships such as "trigger" and "promote" are used as edge labels. Subsequently, while establishing relationship edges, the system retains the causal strength parameters and timestamp attributes of each event, enabling the graph to simultaneously express semantic dependencies and event temporal constraints in its topological structure. After initialization, this forms an agricultural multimodal knowledge foundation framework that supports multi-source knowledge retrieval and logical reasoning.
[0060] This embodiment further performs multimodal mapping operations on the nodes of the initial knowledge graph. By semantically fusing text descriptions, image features, and geospatial information associated with agricultural entities, multimodal features are embedded into a unified node representation. For example, for the node "winter wheat," its text features come from growth period descriptions in agricultural literature, its image features originate from remote sensing images or crop monitoring images, and its spatiotemporal attributes include geographic coordinates and observation time. After fusion mapping, these features are uniformly stored in the node's multimodal vector space, forming a multimodal node set with both semantic and visual information. When new agricultural data arrives, this embodiment performs semantic mapping and fusion calculations on the new data based on agricultural ontology and semantic rules, directly embedding the generated structured knowledge units into the graph. After updating, an expanded agricultural knowledge graph containing newly added semantic entities and event relationships is obtained, achieving dynamic incremental expansion of knowledge content.
[0061] This embodiment, based on an expanded knowledge graph, introduces a confidence-based conflict detection and self-correction mechanism. By calculating the semantic similarity and confidence difference between newly added nodes and existing nodes, it automatically identifies three types of problems: entity overlap, relational contradictions, and temporal conflicts. For example, if a newly inserted event node has a semantic similarity of 0.87 and a temporal conflict intensity of 0.45 with the original event "field irrigation," the system determines it as a potentially redundant node. In this case, this embodiment combines the source weight and temporal confidence in the agricultural evidence template to perform weight adjustment and node merging operations on the conflicting nodes, retaining entities with higher confidence and demoting or removing nodes with lower confidence. Subsequently, a global consistency check is performed on the corrected knowledge graph using causal association and rule-based reasoning mechanisms, and connection state adjustments are performed on edges with low confidence to ensure the logical self-consistency and reasonability of all entities, relations, and events in the knowledge network, ultimately forming an agricultural multimodal knowledge graph that has undergone incremental correction and self-correction.
[0062] Specifically, after incrementally correcting the agricultural knowledge graph, this embodiment first extracts causal relationships, conditional triggering relationships, and outcome dependencies between events from the corrected graph, forming a set of causal associations for agricultural events. This embodiment performs structured parsing of the semantic annotation information of nodes and edges in the knowledge graph, identifying "causal pairs," "trigger pairs," and "outcome pairs" based on the event's temporal attributes, logical order, and semantic type. These correspond to the semantic chains between the event's cause, condition, and result, respectively. For example, "fertilization behavior" triggers "increased soil nitrogen content," and "increased soil nitrogen content" leads to "increased crop growth rate," thus forming a complete causal chain. This embodiment uses joint calculation of causal strength and confidence levels between events to screen out significant causal paths, encapsulating them into a set of event causal associations, providing causal basis for the reasoning calculations of the agricultural knowledge graph.
[0063] This embodiment performs logical reasoning and consistency verification based on agricultural semantic rules and a set of causal relationships. By establishing a knowledge reasoning rule set, it performs logical constraint matching on agricultural entities, relationships, and event nodes to determine whether contradictory relationships, semantic loops, or invalid dependencies exist. For example, when the dependency direction between the "disease control" event and the "pesticide application" event is opposite to known agricultural laws, this embodiment marks the event edge as a logical conflict relationship and outputs a semantic inconsistency warning. Subsequently, this embodiment combines the confidence information in the evidence template to re-evaluate and weight the confidence values of each node and relationship edge in the graph, forming an updated confidence distribution matrix. This matrix stores the changes in the confidence value of each node and edge; for example, an increase from the original value of 0.73 to 0.88 indicates increased node confidence, while a decrease from the original value of 0.65 to 0.42 indicates decreased reliability of the connection.
[0064] After obtaining the updated confidence distribution matrix, this embodiment automatically identifies low-confidence nodes and their associated edges, and optimizes their connection states. Nodes with confidence levels below a set standard (e.g., 0.45) are filtered out using a confidence threshold mechanism, and their connection strength is proportionally reduced or they are reconnected to higher-confidence nodes, thereby reducing the interference of noisy nodes on the overall knowledge structure. After optimization, this embodiment further performs global topology consistency verification and semantic integrity verification. By detecting the node connectivity, causal loop closure, and semantic coverage of the agricultural knowledge graph, it ensures that the adjusted network maintains logical continuity and semantic integrity. The agricultural multimodal knowledge graph output by this process is not only more structurally stable but also achieves self-correction in semantic expression, enabling it to continuously adapt to the dynamic updates and knowledge evolution of agricultural data.
[0065] Furthermore, the specific implementation process of step 800 is as follows:
[0066] This embodiment obtains an updated and self-corrected agricultural multimodal knowledge graph, and then performs evidence fusion reasoning and intelligent recommendation based on this graph to generate agricultural auxiliary decision-making results with confidence level annotations. First, this embodiment extracts event, entity, and relation triples with high causal strength and high confidence from the updated graph to construct a reasoning input subgraph. Combining historical data, monitoring indicators, and expert rule information from the agricultural evidence template, the influence degree of different decision paths is calculated using a causal reasoning mechanism. For example, when there is a significant positive correlation between "decreased soil moisture" and "decreased fertilization efficiency" in the graph and the confidence level is higher than 0.8, this embodiment identifies this causal path as a key influence chain for subsequent decision recommendation analysis.
[0067] In the reasoning process, this embodiment calculates the probability of influence and the degree of conditional dependence between the target event and upstream and downstream events based on agricultural semantic reasoning rules and causal propagation models, generating a causal reasoning matrix. This matrix records the trigger probability and influence intensity of each potential agricultural management event. For example, for the target event of "increased crop yield," the reasoning matrix calculates that the probability of a positive influence from "increased irrigation frequency" is 0.76, and the probability of a negative influence from "insufficient pest and disease control" is 0.64. Based on this, this embodiment further weights the results by considering the credibility of the evidence data, giving greater weight to data from highly reliable sources in the reasoning results, thus ensuring the stability and scientific nature of the reasoning decisions.
[0068] Finally, this embodiment generates multi-dimensional agricultural decision-making support results by integrating the causal inference matrix and confidence distribution. The output results are presented in the form of a set of decision indicators, including recommended measures, influencing factors, and their corresponding confidence intervals. For example, in regional detection, when the model outputs a recommendation of "adjusting irrigation volume by approximately 15%", its confidence level is 0.87, indicating that the decision recommendation has high reliability. This embodiment supports the hierarchical presentation of inference and recommendation results by time, crop type, and production stage, facilitating targeted decision-making and dynamic optimization by agricultural managers, and realizing the intelligent transformation of agricultural production from experience-driven to knowledge- and data-driven.
[0069] This embodiment also provides a multimodal knowledge graph construction system in the agricultural field, including:
[0070] The agricultural ontology construction module is used to collect multi-source agricultural data and analyze the entity types, semantic relationships, and agricultural activity sequences of the multi-source agricultural data in order to construct an agricultural ontology.
[0071] The template generation module is used to construct agricultural semantic rules and evidence templates based on agricultural ontology;
[0072] The fusion module is used to perform semantic matching on the text features and image features of the multi-source agricultural data based on the agricultural ontology and evidence template to obtain preliminary fusion features;
[0073] The optimization module is used to optimize the initial fusion features under ontology constraints to obtain the fusion semantic representation and the corresponding confidence results;
[0074] The structural unit generation module is used to identify agricultural entities, relationships and events based on the fused semantics and the corresponding confidence results, and to perform credibility judgment and conflict resolution based on evidence information to generate structured knowledge units.
[0075] The association network construction module is used to construct causal association networks of events based on structured knowledge units;
[0076] The update module is used to generate an agricultural multimodal knowledge graph based on the event network, and to perform incremental correction and self-correction when new data arrives or conflicts occur, so as to obtain an updated knowledge graph.
[0077] The decision-making module is used to reason and make recommendations based on the updated graph, combined with evidence and causal information, and output agricultural auxiliary decision-making results with confidence.
[0078] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0079] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A method for constructing a multimodal knowledge graph in the agricultural field, characterized in that, include: Collect multi-source agricultural data and analyze the entity types, semantic relationships, and agricultural activity sequences of the multi-source agricultural data in order to construct an agricultural ontology; Construct agricultural semantic rules and evidence templates based on agricultural ontology; Based on the agricultural ontology and evidence template, semantic matching is performed on the text features and image features of the multi-source agricultural data to obtain preliminary fusion features; The initial fusion features are optimized under ontology constraints to obtain the fusion semantic representation and the corresponding confidence results. Agricultural entities, relationships, and events are identified based on the fused semantics and corresponding confidence results, and credibility is determined and conflicts are resolved based on evidence information to generate structured knowledge units. Construct a causal relationship network for events based on structured knowledge units; An agricultural multimodal knowledge graph is generated based on the event network, and incremental corrections and self-corrections are performed when new data arrives or conflicts occur to obtain an updated knowledge graph. Based on the updated graph, reasoning and recommendations are made by combining evidence and causal information, and agricultural auxiliary decision-making results with confidence are output.
2. The method for constructing a multimodal knowledge graph in the agricultural field according to claim 1, characterized in that, The process of collecting and analyzing multi-source agricultural data, including entity types, semantic relationships, and agricultural activity sequences, to construct an agricultural ontology includes: Multi-source agricultural data is collected and its format is standardized and noise is removed to obtain cleaned multi-source agricultural data. The multi-source agricultural data includes text agricultural data and image agricultural data. The text agricultural data includes agricultural literature, agricultural records, operation logs and expert descriptions. The image agricultural data includes crop growth images, agricultural machinery operation images and remote sensing monitoring images. Entities related to agricultural production are identified in the cleaned multi-source agricultural data, and hierarchical labeling is performed according to the category of the entities to obtain a set of agricultural entities, including: crops, farm tools, environmental elements and agricultural activities. Based on domain knowledge and language models, semantic relationships between agricultural entities are extracted to obtain a set of semantic relationships, which includes temporal relationships, causal relationships, and dependency relationships. Based on the stage characteristics of the agricultural production process, a temporal analysis is performed on the set of semantic relations to generate a temporal structure of agricultural activities; Based on the set of agricultural entities, the set of semantic relations, and the temporal structure of agricultural activities, an agricultural ontology is constructed.
3. The method for constructing a multimodal knowledge graph in the agricultural field according to claim 1, characterized in that, The construction of agricultural semantic rules and evidence templates based on agricultural ontology includes: The constraints of entity type, attribute relationship and temporal logic are extracted from the agricultural ontology to establish a semantic constraint set; Based on the set of semantic constraints, agricultural semantic rules reflecting the constraints between entities, behavioral dependencies, and agricultural process logic are formed through a rule generation algorithm. Based on the data types and relational characteristics involved in agricultural semantic rules, evidence fields are defined to identify data sources, collection times, collection methods, and confidence levels. Construct an agricultural evidence template based on the evidence fields.
4. The method for constructing a multimodal knowledge graph in the agricultural field according to claim 1, characterized in that, Based on agricultural ontology and evidence templates, semantic matching is performed on the text features and image features of the multi-source agricultural data to obtain preliminary fusion features, including: Word vectors, entity labels, and contextual semantic features are extracted from textual agricultural data, and target regions, visual feature vectors, and scene labels are extracted from image agricultural data to form textual features and image features, respectively. Based on the entity categories and semantic relationships in the agricultural ontology, semantic mapping is performed on word vectors in text features and visual vectors in image features to obtain semantic mapping results; Based on the semantic mapping results and the confidence weights in the evidence template, the similarity and matching degree between text features and image features are calculated to obtain aligned text features and image features. The aligned text features and image features are weighted and fused to obtain preliminary fused features.
5. The method for constructing a multimodal knowledge graph in the agricultural field according to claim 1, characterized in that, The optimization of the preliminary fusion features under ontology constraints to obtain the fusion semantic representation and corresponding confidence results includes: Construct a set of ontology constraints for feature optimization based on the agricultural ontology; Based on the ontology constraint set, the initial fusion features are adjusted for semantic consistency and feature weights are redistributed to obtain the adjusted fusion features; Multimodal semantic optimization is performed on the adjusted fusion features to generate a fusion semantic representation with a stable semantic structure by maintaining cross-modal semantic consistency and contextual relevance. Based on the data source, time, and credibility weight information in the evidence template, the matching strength and ontology consistency of each feature in the fused semantic representation are evaluated, and the confidence result of each fused semantic unit is calculated and normalized.
6. The method for constructing a multimodal knowledge graph in the agricultural field according to claim 1, characterized in that, The process of identifying agricultural entities, relationships, and events based on fused semantics and corresponding confidence results, and determining credibility and resolving conflicts based on evidence information to generate structured knowledge units includes: Semantic structure analysis is performed on the fused semantic representation to identify the hierarchical relationships and dependency structures of semantic units, thereby obtaining an analyzable agricultural semantic representation model; Based on the entity definition of the agricultural semantic representation model and agricultural ontology, an entity identifier and attribute set are established; Based on agricultural semantic rules, semantic relationships are extracted from the identified entities to form a set of relationships and a set of events; Based on the source, timestamp, and confidence weight in the evidence template, the relevant evidence of the extracted entity, relation set, and event set is matched and weighted to determine the credibility level of each knowledge element; The knowledge elements with confidence conflicts or semantic conflicts are determined based on the credibility level of each knowledge element. By comparing and weighting knowledge elements with conflicting confidence or semantics, a consistent set of agricultural knowledge is obtained. The agricultural knowledge set, relationship set, and event set are grouped together and encapsulated into structured knowledge units to obtain structured knowledge units; The expression for the agricultural semantic representation model is: ; in, For the first A fused representation vector of agricultural semantic units; Embed the optimized agricultural entity features; This is a semantic representation of the relationships between entities; Features of the event context; The confidence weighting factor is obtained from the agricultural evidence template; The first, second, third, and fourth proportional coefficients are used for the system's automatic adaptive learning. It is a nonlinear semantic normalization mapping function.
7. The method for constructing a multimodal knowledge graph in the agricultural field according to claim 1, characterized in that, The expression for the event causal relationship network is: ; in, For the event With the event The strength of the causal relationship between them; and These are the event semantic vectors output by the aforementioned agricultural semantic representation model; This is the causal mapping matrix that the system learns automatically; As a measure of common cause similarity of events; This refers to the time interval or timing difference between events, used to describe the constraints that determine the order of events. and For model adaptive adjustment coefficients; This is the normalized activation function.
8. The method for constructing a multimodal knowledge graph in the agricultural field according to claim 1, characterized in that, The process involves generating an agricultural multimodal knowledge graph based on an event network, and performing incremental corrections and self-corrections when new data arrives or conflicts occur, resulting in an updated knowledge graph, including: Based on the event network, an initial agricultural multimodal knowledge graph is established in the graph database, with agricultural entities as nodes and relationships and events as edges; On the nodes of the initial agricultural multimodal knowledge graph, the text features, image features, and spatiotemporal attributes related to the nodes are multimodal mapped and associated for storage, resulting in a multimodal graph node set with both semantic and visual information; Based on the multimodal graph node set, when new agricultural data arrives, semantic mapping and fusion calculations are performed on the new data according to agricultural ontology and semantic rules to generate new structured knowledge units. The new structured knowledge units are then embedded into the agricultural multimodal knowledge graph to obtain an extended agricultural knowledge graph containing the new knowledge units. For the extended agricultural knowledge graph, the semantic similarity and confidence differences between the newly added nodes and the original nodes and edges are compared to identify entity conflicts, relational conflicts, or temporal conflicts, and a conflict detection result set is obtained. Based on the conflict detection result set, and combined with the evidence template and confidence information, the conflict nodes are updated in weight and merged and adjusted, resulting in an agricultural knowledge graph with incremental correction. Based on the incrementally corrected agricultural knowledge graph, a global consistency check and confidence reassessment are performed using causal association and rule-based reasoning mechanisms. The connection status of low-confidence nodes is automatically adjusted, resulting in an updated and self-correcting agricultural multimodal knowledge graph.
9. A method for constructing a multimodal knowledge graph in the agricultural field according to claim 8, characterized in that, Based on the incrementally corrected agricultural knowledge graph, the process utilizes causal association and rule-based reasoning mechanisms to perform global consistency checks and confidence reassessments, automatically adjusting the connection states of low-confidence nodes, resulting in an updated and self-correcting multimodal agricultural knowledge graph, including: The causal relationships, conditional triggering relationships, and outcome dependencies between events are extracted from the incrementally corrected agricultural knowledge graph to construct a set of causal associations for agricultural events; Based on agricultural semantic rules and the aforementioned causal association set, logical reasoning and consistency verification were performed on the agricultural knowledge graph to identify nodes and relationships with logical contradictions or semantic inconsistencies, and the reasoning and consistency check results were obtained. Based on the reasoning and consistency check results, as well as the confidence information in the evidence template, the confidence values of each node and edge in the agricultural knowledge graph are re-evaluated and weighted, resulting in an updated confidence distribution matrix. Based on the updated confidence distribution matrix, low-confidence nodes and associated edges are automatically identified, and the connection strength of low-confidence nodes and associated edges is attenuated and reconnected, resulting in an agricultural knowledge graph with optimized connection structure. Based on the optimized agricultural knowledge graph with the aforementioned connection structure, global topological consistency verification and semantic integrity verification are performed to obtain an updated and self-correcting agricultural multimodal knowledge graph.
10. A multimodal knowledge graph construction system for the agricultural field, characterized in that, include: The agricultural ontology construction module is used to collect multi-source agricultural data and analyze the entity types, semantic relationships, and agricultural activity sequences of the multi-source agricultural data in order to construct an agricultural ontology. The template generation module is used to construct agricultural semantic rules and evidence templates based on agricultural ontology; The fusion module is used to perform semantic matching on the text features and image features of the multi-source agricultural data based on the agricultural ontology and evidence template to obtain preliminary fusion features; The optimization module is used to optimize the initial fusion features under ontology constraints to obtain the fusion semantic representation and the corresponding confidence results; The structural unit generation module is used to identify agricultural entities, relationships and events based on the fused semantics and the corresponding confidence results, and to perform credibility judgment and conflict resolution based on evidence information to generate structured knowledge units. The association network construction module is used to construct causal association networks of events based on structured knowledge units; The update module is used to generate an agricultural multimodal knowledge graph based on the event network, and to perform incremental correction and self-correction when new data arrives or conflicts occur, so as to obtain an updated knowledge graph. The decision-making module is used to reason and make recommendations based on the updated graph, combined with evidence and causal information, and output agricultural auxiliary decision-making results with confidence.
Citation Information
Cited By
Multi-modal knowledge fusion method and system for carbon emission check data
CN121705468A
Port operation digital sand table construction method and system based on panoramic three-dimensional modeling
CN122089216A
Biological invasion identification method and system based on multi-source data fusion analysis
CN122333289A