Subgraph-driven explainable soil health large language model question and answer method
Patent Information
- Application Number
- CN202610748848.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-28
- Publication Date
- 2026-09-29
AI Technical Summary
[0009]本发明的目的在于提供一种子图驱动可解释的土壤健康大语言模型问答方法,能够支持时空联合查询、多模态融合检索、可解释推理及自动决策生成,有效解决现有技术中土壤健康问答系统缺乏时空动态建模能力、多模态数据融合不充分、问答结果可解释性差、决策支持弱以及无法自动生成可执行决策的问题
[0044]1、增强的时空动态推理能力:通过时间滞后相关性分析和空间邻近边添加,首次在土壤健康知识图谱中实现时空联合查询,能够捕捉指标间的因果关系、空间邻近关系及时间依赖关系。
Smart Images

Figure CN122838532A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the technical field of the intersection of large language modeling and agricultural knowledge engineering, and in particular to a subgraph-driven, interpretable question-answering method for a large language model of soil health. Background Technology
[0002] Soil health is a crucial guarantee for sustainable agricultural development and food security. Traditional soil health assessments rely on laboratory chemical analysis, expert field surveys, and agronomical experience, which suffers from significant drawbacks such as long processing times, high costs, and difficulty in large-scale implementation. In recent years, with the widespread adoption of the Internet of Things (IoT), remote sensing technology, and sensor networks, multimodal soil data—including structured indicators such as pH, organic matter, and heavy metal content, as well as unstructured data such as field images and hyperspectral imagery—has experienced explosive growth, providing a foundation for data-driven soil health management.
[0003] However, existing technologies still have the following significant shortcomings when handling soil health question-answering tasks:
[0004] There is a lack of spatiotemporal dynamic modeling capabilities. Complex causal relationships and time-lag effects exist among soil indicators, and spatial proximity between different geographical locations has a significant impact on soil properties. Existing knowledge graphs or question-answering systems are mostly static structures, unable to support time-series lag correlation queries and spatial proximity inference.
[0005] Multimodal data fusion is insufficient. Existing methods typically process tabular metrics, images, and natural language questions separately, failing to build a unified cross-modal index. This makes it difficult for users to retrieve relevant metric nodes, subgraph structures, and image evidence simultaneously.
[0006] Poor interpretability. The question-and-answer results directly generated by large language models are often "black box" outputs, lacking explanations of the weights of key indicators, demonstrations of causal paths, and citations of visual evidence, making it difficult to meet the needs of agricultural experts and policymakers for credible reasoning.
[0007] The decision generation is not intelligent. The existing system stops at information retrieval and cannot automatically transform the results of questions and answers into actionable decisions such as specific fertilization plans, improvement suggestions, or risk warnings.
[0008] Therefore, there is an urgent need for an intelligent question-answering method for soil health that can integrate spatiotemporal dynamic knowledge graphs, multimodal retrieval-enhanced generation, and high-order interpretable reasoning to solve the above problems. Summary of the Invention
[0009] The purpose of this invention is to provide a subgraph-driven, interpretable large language model question-answering method for soil health, which can support spatiotemporal joint query, multimodal fusion retrieval, interpretable reasoning, and automatic decision generation. It effectively solves the problems of existing soil health question-answering systems, such as lack of spatiotemporal dynamic modeling capabilities, insufficient multimodal data fusion, poor interpretability of question-answering results, weak decision support, and inability to automatically generate executable decisions.
[0010] To achieve the above objectives, the technical solution provided by this invention is: a subgraph-driven, interpretable large language model question-answering method for soil health, comprising the following steps:
[0011] S1: Obtain natural language questions related to soil health and collect multimodal soil data containing indicator tables and images; extract semantic relationships based on the natural language questions and construct a soil health knowledge graph by combining the structured indicators in the indicator tables and images;
[0012] S2: Expand upon the existing soil health knowledge graph by marking the time of nodes on the soil health knowledge graph and adding spatial proximity edges between nodes; through time lag correlation analysis, add time-dependent edges between indicators with significant lag correlations, and finally form a spatiotemporal dynamic knowledge graph that supports spatiotemporal joint queries. The spatiotemporal dynamic knowledge graph consists of indicator nodes, causal relationship edges, spatial proximity edges, and time-dependent edges.
[0013] S3: Divide the spatiotemporal dynamic knowledge graph into subgraphs to obtain multiple subgraphs; vectorize the images in the multimodal data, the indicator nodes in the spatiotemporal dynamic knowledge graph, and the multiple subgraphs respectively to generate image vectors, node vectors, and subgraph vectors, store them in a vector database, and build a cross-modal index based on the stored vectors to support multimodal fusion retrieval;
[0014] S4: Receive user query, vectorize the query statement, perform subgraph-driven retrieval based on the cross-modal index in the vector database, recall the subgraph with the highest semantic similarity to the query and its associated node vectors and image vectors; convert the graph structure of the recalled subgraph into a structured context, input it into the large language model along with the user query, and generate basic question and answer that integrates subgraph information;
[0015] S5: Construct a high-order interpretable prompt template, in which the prompt template embeds the node weights, causal relationship edge types, temporal dependency edge attributes, and spatial proximity edge attributes of the subgraph; inject the basic question and answer, the image vector, the node vector, and the subgraph vector into the prompt template to drive the large language model to perform secondary reasoning and generate interpretable question and answer output containing key indicator interpretation, image evidence citation, and spatiotemporal causal path annotation;
[0016] S6: Based on the interpretable question-and-answer output, automatically generate soil health management decisions, including at least one of the following: fertilization plan, soil improvement suggestions, risk warning, and soil health trend prediction.
[0017] Furthermore, in step S1, the process of constructing the soil health knowledge graph is as follows:
[0018] First, semantic relations are extracted from the natural language problem using named entity recognition and relation extraction techniques. Then, index nodes and initial causal relationship edges are established by combining the structured indicators in the index table and their co-occurrence relationships in the index table. Next, the image is input into a pre-trained visual model to extract image features. Through cross-modal alignment of image features with index nodes, the image is associated as an additional attribute with the corresponding index node, and finally a soil health knowledge graph containing index nodes, causal relationship edges, and image attributes is formed.
[0019] Furthermore, in step S1, the structured indicators in the indicator table include pH value, organic matter content, total nitrogen, available phosphorus, available potassium, cation exchange capacity, pesticide application rate, and heavy metal content.
[0020] Furthermore, in step S2, the time lag correlation analysis specifically involves: targeting any two indicator nodes in the soil health knowledge graph. and We obtain the observed values at different time points, calculate the correlation coefficient and significance level for each lag order, and when there is a lag order that causes the correlation coefficient to exceed the threshold and pass the significance test, we add a time-dependent edge between the two nodes. The calculation formula is as follows:
[0021] ;
[0022] ;
[0023] In the formula, and Representing indicator nodes and In the Observations at each time point; Indicator nodes In comparison time morning Observations at a given time point; , This represents the average of the observations of the two indicator nodes across all observation time points. For observation time series; Indicator Node In lag order Next indicator node The correlation coefficient; The correlation coefficient represents the probability value for significance assessment; the smaller the value, the more reliable the correlation. The maximum lag order, The correlation threshold, The significance level; This is an edge existence indicator function; a value of 1 indicates the existence of a time-dependent edge, and a value of 0 indicates the absence of a time-dependent edge. When adding, the direction is The lag order is Time-dependent edges.
[0024] Furthermore, in step S3, a community detection algorithm is used to subdivide the spatiotemporal dynamic knowledge graph. This community detection algorithm aims to maximize modularity by iteratively merging nodes or edges to divide the spatiotemporal dynamic knowledge graph into several subgraphs with tightly connected internal connections and sparse external connections, thus maximizing modularity. A larger value indicates a tighter internal connection and a sparser connection between communities, resulting in a better partitioning effect. The formula is as follows:
[0025] ;
[0026] In the formula, and For any two indicator nodes in the spatiotemporal dynamic knowledge graph; Indicator Node and The weight of the edges between them is 0 if there is no connecting edge, and 0 if there is a connecting edge. The weight is equal to the weight of the edge, which is determined according to the type of edge: the weight of a causal edge is set to 1 by default, the weight of a spatially adjacent edge is set to the reciprocal of the spatial distance between the two nodes, and the weight of a time-dependent edge is set to the absolute value of the corresponding lag correlation coefficient. and For the corresponding indicator nodes and The sum of the weights of all connected edges; The total weight of all edges in the spatiotemporal dynamic knowledge graph; and Indicates the corresponding indicator node and The community to which it was assigned; It is a conditional function, when and The value is 1 if they are in the same community, and 0 otherwise. Each community in the community detection results constitutes a subgraph, and each subgraph contains a complete structure of indicator nodes, causal relationship edges, spatial proximity edges, and temporal dependency edges.
[0027] Further, in step S3, the specific process of vectorization is as follows: each indicator node in the spatiotemporal dynamic knowledge graph is encoded to obtain a node vector; each subgraph is aggregated to obtain a subgraph vector; global features are extracted from the images in the multimodal data to obtain image vectors; based on the above node vectors, subgraph vectors, and image vectors, a cross-modal index is constructed and associatedly stored by cross-modal alignment of node vectors, subgraph vectors, and image vectors, expressed by the formula:
[0028] ;
[0029] ;
[0030] ;
[0031] In the formula, This represents an index node in a spatiotemporal dynamic knowledge graph. GNN stands for Graph Neural Network encoding operation, used to transform the connection information of the index node into a numerical vector. Indicator Node The node vector obtained after encoding by a graph neural network; This represents a subgraph. READOUT is an aggregation operation that combines multiple vectors into a single vector. For subgraph The subgraph vector obtained by summing the vectors of all its internal nodes; For images in multimodal data, ViT is the encoding operation of the Visual Transformer model, used to extract overall image features. For image The corresponding image vector;
[0032] The cross-modal index is constructed by aligning it in a shared embedding space through multimodal contrastive learning. , and The distribution of the vectors makes the cosine similarity between node vectors, subgraph vectors and image vectors that describe similar soil health semantics higher; the cross-modal index stores each subgraph vector, node vectors within the subgraph and image vectors associated with the subgraph in the vector database to support retrieving data of other modalities using text, images or subgraphs as queries.
[0033] Furthermore, in step S4, the specific process of the subgraph-driven retrieval is as follows: [The user query is then processed / retrieved]. The query vector is vectorized using a text encoder, which is a model that converts natural language statements into numerical vectors. The similarity between the query vector and the vectors of each subgraph is calculated in a vector database based on a cross-modal index.
[0034] ;
[0035] In the formula, , , are adjustable weight parameters, which are used to balance the contribution of the whole subgraph, key nodes in the subgraph and associated images to the similarity; represents the cosine similarity between two vectors, the closer the cosine similarity value is to 1, the more similar the two vectors are; is the user query obtained by a text encoder a query vector after vectorization, is a key node vector in the corresponding subgraph; represents the user query and the subgraph similarity, sort the similarity from high to low, recall the most relevant subgraph and its associated node vectors and image vectors; convert the graph structure of the recalled subgraph into structured context in natural language form, splice it with the user query and input it into a large language model to generate a basic question and answer integrating subgraph information.
[0036] Further, in step S5, a high-order interpretable prompt template is constructed , the construction method is:
[0037] ;
[0038] In the formula, the prompt template is organized in the following format: is the basic question and answer content generated in step S4, represents the topological structure and edge types of the subgraph, represents a node importance weight matrix, represents a set of causal edge types, represents a set of time-dependent edges and their lag properties, represents a set of spatial proximity edges; is an image vector set, is a node vector set, is a subgraph vector set; the secondary reasoning is performed by inputting the prompt template into a large language model, so that the large language model can simultaneously perceive graph structure constraints, spatio-temporal properties and multimodal semantics, and generate interpretable question and answer output containing key indicator interpretation, image evidence citation and spatio-temporal causal path annotation .
[0039] Furthermore, in step S5, the node weights are calculated on the subgraph using a graph attention mechanism, and the index nodes... Node weights The soil indicators most influential on user queries are determined by attention scores and used to highlight them in the prompt template; the causal relationship edge types include facilitator, inhibitor, and converter, which are injected into the prompt template as relational tags in a serialized manner; the time-dependent edge attributes include lag order and correlation coefficient; and the spatial proximity edge attributes include proximity distance and relative positional relationship.
[0040] Furthermore, in step S6, the specific process of automatically generating soil health management decisions is as follows: outputting the interpretable question-and-answer generated in step S5. User query Combined into decision prompts, input into a large language model Generate structured decision text :
[0041] ;
[0042] This structured decision text It should include at least one of the following: fertilization plan, soil improvement recommendations, risk warning level, and soil health trend prediction, along with decision-making criteria based on spatiotemporal causal paths and node weights; when a risk warning is triggered, this structured decision text... The information further includes warning thresholds and corresponding time and spatial range information.
[0043] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0044] 1. Enhanced spatiotemporal dynamic reasoning ability: Through time lag correlation analysis and the addition of spatial proximity edges, spatiotemporal joint query is realized for the first time in the soil health knowledge graph, which can capture the causal relationship, spatial proximity relationship and time dependency relationship between indicators.
[0045] 2. Highly efficient multimodal fusion retrieval: Employing subgraph partitioning and cross-modal indexing techniques, node vectors, subgraph vectors, and image vectors are uniformly mapped to a shared embedding space, supporting subgraph-driven multimodal retrieval and significantly improving recall relevance and retrieval efficiency.
[0046] 3. High-order interpretability: Node weights are calculated through graph attention mechanism, and causal edge types and spatiotemporal attributes are explicitly embedded in the prompt template, so that the secondary reasoning results of the large language model have a traceable and verifiable chain of evidence.
[0047] 4. Intelligent decision-making closed loop: An automated closed loop is formed from question and answer to decision-making, directly outputting fertilization plans, improvement suggestions, risk warnings and trend predictions to serve the actual agricultural production.
[0048] In summary, this invention can effectively solve the problems of existing soil health question-and-answer systems, such as lack of spatiotemporal modeling, insufficient multimodal fusion, poor interpretability, and weak decision support. Attached Figure Description
[0049] Figure 1 This is a flowchart illustrating the overall process framework of the method of this invention. In the diagram, S1 is the construction of a soil health knowledge graph, S2 is the expansion of a spatiotemporal dynamic knowledge graph, S3 is subgraph partitioning and multimodal vectorization, S4 is subgraph-driven retrieval and basic question-and-answer generation, S5 is a high-order interpretable prompt template and secondary reasoning, and S6 is the generation of soil health management decisions. Natural language relationships, indicator tables, and image information are the data sources for S1; time, space, and causality are the expansion targets of S2; the vector database is the output storage for S3; user query input, retrieval, and then structured context are the inputs for S4; key indicators, images, and spatiotemporal causality are the auxiliary information for S5; and decisions, solutions, suggestions, warnings, and trends are the outputs of S6.
[0050] Figure 2 This is a schematic diagram of the spatiotemporal causal network among soil health indicators. In the diagram, pH, organic matter, total nitrogen, available phosphorus, available potassium, heavy metals, cation exchange capacity, and pesticide application rate are nodes of soil health indicators. Solid arrows indicate the promoting or inhibiting relationship between indicators. Dashed arrows indicate spatiotemporal constraints, where lag=3 months from organic matter to heavy metals indicates a time lag, lag=3 months from heavy metals to cation exchange capacity indicates a time lag, and the nearest 500m from cation exchange capacity to pH indicates a spatial proximity constraint.
[0051] Figure 3 This diagram illustrates the subgraph partitioning and cross-modal indexing construction. In the diagram, the complete spatiotemporal dynamic knowledge graph is the input graph, and subgraphs A, B, and C are the partitioned subgraphs. GNN encoding is used to generate node vectors, READOUT aggregation is used to generate subgraph vectors, and ViT encoding is used to generate image vectors. The vector database is a repository for multimodal vectors, where cross-modal indexing represents the indexing method of multimodal vectors, and shared embedding space and cosine similarity alignment represent the alignment strategy between vectors. Detailed Implementation
[0052] The present invention will be further described in detail below with reference to the embodiments and accompanying drawings, but the embodiments of the present invention are not limited thereto.
[0053] like Figure 1 As shown, this embodiment discloses a subgraph-driven, interpretable question-answering method for a large language model of soil health. Through techniques such as constructing a spatiotemporal dynamic knowledge graph, cross-modal indexing, subgraph-driven retrieval, and high-order interpretable prompt templates, it achieves an intelligent closed loop from multimodal data to interpretable decision-making. The specific implementation includes the following steps:
[0054] 1) Constructing a knowledge graph of soil health
[0055] First, data sources need to be acquired. On one hand, natural language questions related to soil health are needed, such as "Why has the yield of this land declined year after year?" and "How can acidified soil be improved?". On the other hand, multimodal soil data, including indicator tables and images, needs to be collected simultaneously. The indicator tables contain structured soil indicators, including but not limited to: pH value, organic matter content, total nitrogen, available phosphorus, available potassium, cation exchange capacity, pesticide application rate, and heavy metal content. Images include field soil profile images, crop growth images, or remote sensing images.
[0056] Based on the above data sources, we began to construct a soil health knowledge graph. The specific construction process is as follows:
[0057] Semantic Relation Extraction: Utilizing named entity recognition and relation extraction techniques, semantic relations are extracted from the acquired natural language questions. For example, a semantic relationship of "influence" or "depends on" can be identified between the entity "crop yield" and the entity "soil organic matter" in the question.
[0058] Establishing Indicator Nodes and Causal Relationship Edges: Based on the structured indicators in the indicator table, each specific soil indicator is established as an independent node. Simultaneously, based on the co-occurrence relationships of the indicators in the table and the known soil science mechanisms between the indicators, initial causal relationship edges are established between the indicator nodes. For example, an "Influence" edge is established between the "pH value" node and the "available phosphorus" node, because soil pH directly affects the availability of phosphorus.
[0059] Image Feature Association: The acquired images are input into a pre-trained visual Transformer model to extract image features. Subsequently, through cross-modal alignment techniques, the extracted image feature vectors are associated with the most semantically relevant indicator nodes in the knowledge graph. For example, an image showing soil compaction and rust spots will have its feature vector associated with indicator nodes such as "soil structure" or "redox potential". After association, the image is stored in the knowledge graph as an additional attribute of the corresponding indicator node.
[0060] 2) Constructing a spatiotemporal dynamic knowledge graph
[0061] Based on the constructed static soil health knowledge graph, spatiotemporal dimensions are expanded to form a spatiotemporal dynamic knowledge graph that supports spatiotemporal joint queries. For example... Figure 2 As shown, the expansion process mainly consists of three parts:
[0062] 2.1) Labeling the time attribute: Label each indicator node on the soil health knowledge graph with the timestamp of its data generation or collection. If an indicator node represents "the organic matter content of plot A in June 2023", then the indicator node has a clear time attribute;
[0063] 2.2) Adding Spatial Proximity Edges: Based on the geographical location information of the data collection points, calculate the spatial distance between any two indicator nodes representing the land parcels. When the spatial distance between two indicator nodes is less than a preset proximity distance threshold, add a "spatial proximity edge" between these two indicator nodes to represent their geographical proximity. The attributes of this edge can include proximity distance and relative positional relationship;
[0064] 2.3) Adding time-dependent edges: This is the core of constructing a spatiotemporal dynamic knowledge graph. Through time-lag correlation analysis, time-dependent edges are added between indicators with significant lag correlations. The specific analysis process is as follows:
[0065] For any two indicator nodes in the soil health knowledge graph and This allows us to obtain their respective historical observation sequences at different points in time. For example, obtaining the observation time series over the past 36 months. ,node Such as organic matter content and indicator nodes Such as monthly observation data on crop yield;
[0066] Set a maximum lag order ,For example Months. Then, for each possible lag order... Calculate indicator nodes Lag Lower-order pairs of nodes The Pearson correlation coefficient is calculated using the following formula:
[0067] ;
[0068] In the formula, and Representing indicator nodes and In the Observations at each time point; Indicator nodes In comparison time morning Observations at a given time point; , This represents the average of the observations of the two indicator nodes across all observation time points. For observation time series; Indicator Node In lag order Next indicator node The correlation coefficient;
[0069] For each calculated correlation coefficient Perform a significance test to obtain the corresponding significance probability value. ; Traverse all lag orders Check if there is at least one This makes the absolute value of its correlation coefficient... Exceeding the preset correlation threshold , and its The value is less than the preset significance level. If the condition is met, then the edge has an indicator function. The value is 1:
[0070] ;
[0071] In the formula, This is an edge existence indicator function; a value of 1 indicates the existence of a time-dependent edge, and a value of 0 indicates the absence of a time-dependent edge. When adding, the direction is The lag order is Time-dependent edges, meaning " Past values affect "The current value of the edge." The properties of this edge include the lag order that makes it valid. and the corresponding correlation coefficient .
[0072] For example, this analysis might reveal "fertilizer application rate". "Lag of 6 months" Time and "soil organic matter content" If a significant positive correlation is found, a time-dependent edge with a lag order of 6 will be added between the two, pointing from "fertilizer application amount" to "organic matter content". This accurately captures the time lag effect that organic fertilizer takes several months to fully decompose and increase organic matter after being applied to the soil.
[0073] Ultimately, the spatiotemporal dynamic knowledge graph that supports spatiotemporal joint queries is formed by indicator nodes, causal relationship edges, spatial proximity edges, and temporal dependency edges.
[0074] 3) Subgraph partitioning, vectorization, and cross-modal index construction
[0075] To improve the efficiency and accuracy of subsequent retrievals, this step processes the spatiotemporal dynamic knowledge graph, such as... Figure 3 As shown, it mainly includes three stages: subgraph partitioning, vectorization processing, and cross-modal index construction.
[0076] 3.1) Subgraph Partitioning: The spatiotemporal dynamic knowledge graph is partitioned into subgraphs using a community detection algorithm. This algorithm aims to maximize modularity, and the formula for calculating modularity Q is:
[0077] ;
[0078] In the formula, and For any two indicator nodes in the knowledge graph; Indicator Node and The weight of the edges between them is 0 if there is no connecting edge; if there is a connecting edge, then... The weight is equal to the weight of the edge, which is determined according to the type of edge: the weight of a causal edge is set to 1 by default, the weight of a spatially adjacent edge is set to the reciprocal of the spatial distance between the two nodes, and the weight of a time-dependent edge is set to the absolute value of the corresponding lag correlation coefficient. and For the corresponding node and The sum of the weights of all connected edges; The total weight of all edges in the spatiotemporal dynamic knowledge graph; and Indicates the corresponding node and The community to which it was assigned; It is a conditional function, when and The value is 1 if the nodes or edges belong to the same community, and 0 otherwise; the algorithm iteratively merges nodes or edges to achieve this. The value continues to increase.
[0079] The larger the value, the denser the connections within the divided communities and the sparser the connections between communities, resulting in a better partitioning effect. Each community in the community detection results constitutes a subgraph. Ultimately, the entire spatiotemporal dynamic knowledge graph is divided into multiple subgraphs, each of which fully preserves the indicator nodes, causal relationship edges, spatial proximity edges, and temporal dependency edges within that community.
[0080] 3.2) Vectorization Processing: The images in the multimodal data, each indicator node in the spatiotemporal dynamic knowledge graph, and the multiple subgraphs are each vectorized to generate corresponding numerical vectors.
[0081] Node vectorization: For each index node in the graph The index node is encoded using a graph neural network (GNN). The encoding operation considers not only the attributes of the index node itself but also aggregates information from its neighboring index nodes, thereby generating an index node vector containing structural information. The formula is: .
[0082] Subgraph vectorization: For each subgraph obtained from the partitioning By using an aggregation operation, the READOUT function, such as average pooling or summation pooling, the vectors of all nodes within the subgraph are aggregated, thereby generating a subgraph vector that can represent the semantics of the entire subgraph. The formula is: .
[0083] Image vectorization: For images in multimodal data The visual Transformer model is used to extract its global features and generate image vectors. The formula is: .
[0084] 3.3) Constructing a Cross-Modal Index: To enable data of different types to be retrieved from each other, a cross-modal index needs to be constructed. The construction method employs multimodal contrastive learning technology. During training, the node vectors generated above are used... Subgraph vectors With image vectors The data is mapped to a shared embedding space, and their distributions are aligned. The goal is to ensure that node vectors, subgraph vectors, and image vectors describing similar soil health semantics have higher cosine similarity among vectors in this space. After alignment, the data is stored in a vector database, specifically organized as follows: each subgraph vector is stored in association with the others in the vector database. All node vectors within this subgraph and the image vectors related to the content of this subgraph. This associated storage structure is called a cross-modal index. Based on this index, users can input any form of text, image, or subgraph information as a query to retrieve data from other modalities that are semantically related to it.
[0085] 4) Subgraph-driven retrieval and basic question-answering generation
[0086] This step involves receiving and initially responding to user queries.
[0087] 4.1) Query Vectorization: Receiving user queries This is natural language text. Using a text encoder, which is a model that converts natural language statements into numerical vectors, ... Convert to query vector .
[0088] 4.2) Subgraph-driven retrieval: In a vector database, the query vector is calculated based on the constructed cross-modal index. With all subgraphs in the database The similarity is calculated. To more comprehensively evaluate similarity, this method designs a comprehensive similarity score. The calculation formula is as follows:
[0089] ;
[0090] In the formula, , , These are adjustable weight parameters used to balance the contributions of the subgraph as a whole, key nodes within the subgraph, and related images to the similarity score. The cosine similarity between two vectors is represented by a value closer to 1, indicating greater similarity. The values are sorted from high to low, and the one or more subgraphs with the highest recall score, i.e. the highest similarity, along with their associated node vectors and image vectors, are retrieved together.
[0091] 4.3) Generate basic question answering: Transform the graph structure of the highest-scoring subgraph into a structured context in human-understandable natural language. For example, transform the causal relationship edges in the subgraph into a structured context. The sentence is translated to "The low pH in this region inhibits the activity of available phosphorus." This structured context is then compared with the user's original query. The two parts are concatenated to form an enhanced prompt word, which is then input into the large language model. The large language model uses this prompt word, which incorporates subgraph information, for reasoning and text generation, resulting in a richer and more well-supported basic question-and-answer framework that integrates subgraph information. .
[0092] 5) Higher-order interpretable reasoning and question-answering output
[0093] To address the issue of "black box" output from large language models, this step constructs a high-order interpretable prompt template to drive the large language model to perform secondary reasoning, thereby generating interpretable question-and-answer responses with a chain of evidence.
[0094] 5.1) Calculating Node Weights: To highlight key points in the explanation, it is necessary to quantify the importance of each indicator node in the subgraph to the user query. A graph attention mechanism is used to perform calculations on the recalled subgraph, weighting each indicator node... Calculate an attention score, which serves as the indicator node. Node weights This ultimately forms the importance weight matrix of the indicator nodes. A node with a higher weight is more crucial to answering user questions.
[0095] 5.2) Constructing a higher-order interpretable hint template: Construct a hint template P with a structured parameter domain, the construction formula of which is expressed as:
[0096] ;
[0097] This prompt template P is organized according to the following specific format and content:
[0098] Enter the generated basic question and answer content.
[0099] Enter the topology of the recall subgraph and the type of edges to indicate how the model nodes are connected.
[0100] Enter the node importance weight matrix, highlighting the most important indicators in the form of "[Key Indicator] XXX (Weight 0.8)".
[0101] : Fill in the set of causal relationship edge types. These types specifically include "promote", "inhibit", and "transform", and are serialized and injected into the template in the form of "relationship label: [node1] promote [node2]".
[0102] : Fill in the set of time-dependent edges and their lag attributes. Attributes include lag order and correlation coefficient, and are specified in the template as "Time Dependency: [Node 3] in lag [ The injection is performed in the form of "[Node 4](correlation coefficient: [value])".
[0103] : Fill in the set of spatially adjacent edges. Attributes include proximity distance and relative positional relationship, which are injected into the template in the form of "Spatial relationship: [Node 5] is adjacent to [Node 6] (distance: [value] meters)".
[0104] , , : Fill in the sets of associated image vectors, node vectors and subgraph vectors respectively to provide underlying multimodal semantic information.
[0105] 5.3) Generate interpretable question-answering output: This involves constructing a high-order interpretable hint template that embeds rich graph structures, spatiotemporal attributes, and multimodal semantic information. Inputting a large language model drives the model to perform secondary reasoning. After perceiving these graph structure constraints, the model generates answers that are no longer generalities, but rather interpretable answers containing cited evidence. The final interpretable question-and-answer output includes: interpretation of key indicators, such as "the most critical factor in this analysis is soil pH value", citation of image evidence, such as "the evidence comes from images collected in March 2023, showing obvious signs of acidification in the area", and spatiotemporal causal path annotation, such as "according to the knowledge graph, long-term excessive application of chemical fertilizers [cause] → (6-month lag) → soil acidification [result] → (spatial diffusion) → risk of eutrophication in surrounding water bodies [warning]".
[0106] 6) Soil health management decision generation
[0107] This step completes the closed loop from "question and answer" to "decision-making." The generated interpretable question-and-answer output is then completed. Compared with the user's original query These elements are combined to construct a decision prompt. This decision prompt is then input into a large language model to generate structured decision text. The process is represented as follows:
[0108] ;
[0109] The large language model automatically generates practical soil health management decisions based on reasoning and evidence in interpretable question-and-answer outputs. This structured decision text... It must contain at least one of the following:
[0110] Fertilization plan: For example, "It is recommended to apply 1,500 kg of well-rotted cow manure per mu (unit of land area) along with 20 kg of compound fertilizer as base fertilizer before the next planting season."
[0111] Soil improvement recommendations: For example, "Based on the trend of soil acidification, it is recommended to apply lime for improvement, with a target pH of 6.5."
[0112] Risk Warning: Triggered when the reasoning result reaches a preset threshold. The decision text will not only issue a warning, but also include the warning threshold and the corresponding time and spatial range information, such as "[Risk Warning] Warning Level: High. The cadmium content in the soil of the southeast area of this plot (spatial range) has been rising continuously over the past 12 months (time range), approaching the screening value of the national soil environmental quality standard (warning threshold). It is recommended to take measures to apply passivation materials within the next 6 months (time range)."
[0113] Soil health trend forecasts: For example, "Under the current management measures, the soil organic matter content is expected to increase steadily at a rate of 0.1% per year over the next 3 years."
[0114] The resulting structured decision text not only provides specific action recommendations, but also includes decision-making criteria based on spatiotemporal causal paths and node weights, making the decision-making process transparent and credible, and directly serving agricultural production practices.
[0115] The above embodiments are preferred embodiments of the present invention, but the embodiments of the present invention are not limited to the above embodiments. Any changes, modifications, substitutions, combinations, or simplifications made without departing from the spirit and principle of the present invention shall be considered equivalent substitutions and shall be included within the protection scope of the present invention.
Claims
1. A subgraph-driven, interpretable large language model for soil health question answering, characterized in that, Includes the following steps: S1: Obtain natural language questions related to soil health and collect multimodal soil data containing indicator tables and images; extract semantic relationships based on the natural language questions and construct a soil health knowledge graph by combining the structured indicators in the indicator tables and images; S2: Expand upon the existing soil health knowledge graph by marking the time of nodes on the soil health knowledge graph and adding spatial proximity edges between nodes; through time lag correlation analysis, add time-dependent edges between indicators with significant lag correlations, and finally form a spatiotemporal dynamic knowledge graph that supports spatiotemporal joint queries. The spatiotemporal dynamic knowledge graph consists of indicator nodes, causal relationship edges, spatial proximity edges, and time-dependent edges. S3: Divide the spatiotemporal dynamic knowledge graph into subgraphs to obtain multiple subgraphs; vectorize the images in the multimodal data, the indicator nodes in the spatiotemporal dynamic knowledge graph, and the multiple subgraphs respectively to generate image vectors, node vectors, and subgraph vectors, store them in a vector database, and build a cross-modal index based on the stored vectors to support multimodal fusion retrieval; S4: Receive user query, vectorize the query statement, perform subgraph-driven retrieval based on the cross-modal index in the vector database, recall the subgraph with the highest semantic similarity to the query and its associated node vectors and image vectors; convert the graph structure of the recalled subgraph into a structured context, input it into the large language model along with the user query, and generate basic question and answer that integrates subgraph information; S5: Construct a high-order interpretable hint template, wherein the hint template embeds the node weights, causal edge types, temporal dependency edge attributes, and spatial proximity edge attributes of the subgraph; The basic question and answer, the image vector, the node vector, and the subgraph vector are all injected into the prompt template to drive the large language model to perform secondary reasoning and generate an interpretable question and answer output that includes interpretation of key indicators, reference of image evidence, and spatiotemporal causal path annotation. S6: Based on the interpretable question-and-answer output, automatically generate soil health management decisions, including at least one of the following: fertilization plan, soil improvement suggestions, risk warning, and soil health trend prediction.
2. The subgraph-driven interpretable large language model question-answering method for soil health according to claim 1, characterized in that, In step S1, the process of constructing the soil health knowledge graph is as follows: First, semantic relations are extracted from the natural language problem using named entity recognition and relation extraction techniques. Then, index nodes and initial causal relationship edges are established by combining the structured indicators in the index table and their co-occurrence relationships in the index table. Next, the image is input into a pre-trained visual model to extract image features. Through cross-modal alignment of image features with index nodes, the image is associated as an additional attribute with the corresponding index node, and finally a soil health knowledge graph containing index nodes, causal relationship edges, and image attributes is formed.
3. The subgraph-driven, interpretable large language model question-answering method for soil health according to claim 1, characterized in that, In step S1, the structured indicators in the indicator table include pH value, organic matter content, total nitrogen, available phosphorus, available potassium, cation exchange capacity, pesticide application rate, and heavy metal content.
4. The subgraph-driven, interpretable large language model question-answering method for soil health according to claim 1, characterized in that, In step S2, the time lag correlation analysis specifically involves: targeting any two indicator nodes in the soil health knowledge graph. and We obtain the observed values at different time points, calculate the correlation coefficient and significance level for each lag order, and when there is a lag order that causes the correlation coefficient to exceed the threshold and pass the significance test, we add a time-dependent edge between the two nodes. The calculation formula is as follows: ; ; In the formula, and Representing indicator nodes and In the Observations at each time point; Indicator nodes In comparison time morning Observations at a given time point; , This represents the average of the observations of the two indicator nodes across all observation time points. For observation time series; Indicator Node In lag order Next indicator node The correlation coefficient; The correlation coefficient represents the probability value for significance assessment; the smaller the value, the more reliable the correlation. The maximum lag order, This is the correlation threshold. The significance level; This is an edge existence indicator function; a value of 1 indicates the existence of a time-dependent edge, and a value of 0 indicates the absence of a time-dependent edge. When adding, the direction is The lag order is Time-dependent edges.
5. The subgraph-driven, interpretable large language model question-answering method for soil health according to claim 1, characterized in that, In step S3, a community detection algorithm is used to partition the spatiotemporal dynamic knowledge graph into subgraphs. This algorithm aims to maximize modularity by iteratively merging nodes or edges to divide the spatiotemporal dynamic knowledge graph into several subgraphs with tightly connected internal connections and sparse external connections, thus maximizing modularity. A larger value indicates a tighter internal connection and a sparser connection between communities, resulting in a better partitioning effect. The formula is as follows: ; In the formula, and For any two indicator nodes in the spatiotemporal dynamic knowledge graph; Indicator Node and The weight of the edges between them is 0 if there is no connecting edge, and 0 if there is a connecting edge. The weight is equal to the weight of the edge, which is determined according to the type of edge: the weight of a causal edge is set to 1 by default, the weight of a spatially adjacent edge is set to the reciprocal of the spatial distance between the two nodes, and the weight of a time-dependent edge is set to the absolute value of the corresponding lag correlation coefficient. and For the corresponding indicator nodes and The sum of the weights of all connected edges; The total weight of all edges in the spatiotemporal dynamic knowledge graph; and Indicates the corresponding indicator node and The community to which it was assigned; It is a conditional function, when and The value is 1 if they are in the same community, and 0 otherwise. Each community in the community detection results constitutes a subgraph, and each subgraph contains a complete structure of indicator nodes, causal relationship edges, spatial proximity edges, and temporal dependency edges.
6. The subgraph-driven, interpretable large language model question-answering method for soil health according to claim 5, characterized in that, In step S3, the specific process of vectorization is as follows: each indicator node in the spatiotemporal dynamic knowledge graph is encoded to obtain a node vector; each subgraph is aggregated to obtain a subgraph vector; global features are extracted from the images in the multimodal data to obtain image vectors; based on the above node vectors, subgraph vectors, and image vectors, a cross-modal index is constructed and associatedly stored by cross-modal alignment of node vectors, subgraph vectors, and image vectors, expressed by the formula: ; ; ; In the formula, This represents an index node in a spatiotemporal dynamic knowledge graph. GNN stands for Graph Neural Network encoding operation, used to transform the connection information of the index node into a numerical vector. Indicator Node The node vector obtained after encoding by a graph neural network; This represents a subgraph. READOUT is an aggregation operation that combines multiple vectors into a single vector. For subgraph The subgraph vector obtained by summing the vectors of all its internal nodes; For images in multimodal data, ViT is the encoding operation of the Visual Transformer model, used to extract overall image features. For image The corresponding image vector; The cross-modal index is constructed by aligning it in a shared embedding space through multimodal contrastive learning. , and The distribution of these vectors makes the cosine similarity between node vectors, subgraph vectors, and image vectors that describe similar soil health semantics higher. Cross-modal indexes store each subgraph vector, the node vectors within the subgraph, and the image vectors associated with that subgraph in a vector database, to support retrieving data from other modalities using text, images, or any form of subgraph as a query.
7. The subgraph-driven, interpretable large language model question-answering method for soil health according to claim 6, characterized in that, In step S4, the specific process of the subgraph-driven retrieval is as follows: The user query... The query vector is vectorized using a text encoder, which is a model that converts natural language statements into numerical vectors. The similarity between the query vector and the vectors of each subgraph is calculated in a vector database based on a cross-modal index. ; In the formula, , , are adjustable weight parameters, which are used to balance the contribution of the overall subgraph, key nodes in the subgraph and associated images to the similarity; represents the cosine similarity between two vectors, and the closer the value of the cosine similarity is to 1, the more similar the two vectors are; is a query vector obtained after vectorizing the user query by a text encoder, is a key node vector in the corresponding subgraph; represents the similarity between the user query and the subgraph is sorted from high similarity to low similarity, and the most relevant subgraph and its associated node vector and image vector are recalled; the graph structure of the recalled subgraph is converted into structured context in the form of natural language, which is spliced with the user query and then input into a large language model to generate a basic question answering integrating subgraph information.
8. The subgraph-driven interpretable large language model question-answering method for soil health according to claim 7, characterized in that, In step S5, a higher-order interpretable prompt template is constructed. The construction method is as follows: ; In the formula, the prompt template Organize in the following format: The basic question and answer content generated in step S4, This indicates the topology of the subgraph and the type of its edges. This represents the node importance weight matrix. A set representing the type of causal edge. This represents the set of time-dependent edges and their lag properties. Represents the set of spatially adjacent edges; Image vector The set, node vector The set, Subgraph vector The set; the secondary reasoning is achieved by using the prompt template Inputting the large language model enables it to simultaneously perceive graph structure constraints, spatiotemporal attributes, and multimodal semantics, generating interpretable question-and-answer outputs that include interpretation of key indicators, citation of image evidence, and spatiotemporal causal path annotation. .
9. The subgraph-driven, interpretable large language model question-answering method for soil health according to claim 8, characterized in that, In step S5, the node weights are calculated on the subgraph using a graph attention mechanism, and the index nodes... Node weights The soil indicators most influential on user queries are determined by attention scores and used to highlight them in the prompt template; the causal relationship edge types include facilitator, inhibitor, and converter, which are injected into the prompt template as relational tags in a serialized manner; the time-dependent edge attributes include lag order and correlation coefficient; and the spatial proximity edge attributes include proximity distance and relative positional relationship.
10. The subgraph-driven, interpretable large language model question-answering method for soil health according to claim 9, characterized in that, In step S6, the specific process of automatically generating soil health management decisions is as follows: outputting the interpretable question and answer generated in step S5. User query Combined into decision prompts, input into a large language model Generate structured decision text : ; This structured decision text It should include at least one of the following: fertilization plan, soil improvement recommendations, risk warning level, and soil health trend prediction, along with decision-making criteria based on spatiotemporal causal paths and node weights; when a risk warning is triggered, this structured decision text... The information further includes warning thresholds and corresponding time and spatial range information.