A tea garden risk prediction method and system based on multi-modal data

By using high-dimensional survival analysis and dynamic causal discovery of multimodal data, a tea garden risk prediction model was constructed, which solved the problem of insufficient causal relationship modeling in tea garden risk prediction and achieved efficient and interpretable risk prediction and decision support.

CN122334933APending Publication Date: 2026-07-03HANGZHOU XIANGCHAN TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HANGZHOU XIANGCHAN TECHNOLOGY CO LTD
Filing Date
2026-02-12
Publication Date
2026-07-03

AI Technical Summary

Technical Problem

Existing tea garden risk prediction methods lack high-reliability feature denoising and causal relationship modeling capabilities when processing multimodal data, resulting in high false alarm rates, lagging time series predictions, and weak mechanism explanations. They are unable to adapt to the microclimate variations and sudden disaster scenarios in complex mountain tea gardens.

Method used

By employing high-dimensional survival analysis based on multimodal data, contribution screening, spatiotemporal feature fusion, dynamic causal discovery, and a hybrid expert architecture, a modal causal graph is constructed and cross-modal feature fusion is performed to generate a tea garden risk report.

Benefits of technology

It significantly improves the stability and generalization ability of tea garden risk prediction, provides high-confidence and highly interpretable risk prediction results, and can characterize the differences and spatiotemporal characteristics of the impact of environmental factors on risk, supporting reliable decision optimization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122334933A_ABST
    Figure CN122334933A_ABST
Patent Text Reader

Abstract

The application provides a tea garden risk prediction method and system based on multi-modal data, comprising: based on risk correlation constraints, high-dimensional survival analysis and contribution degree filtering and noise removal are performed on a tea garden multi-modal data set, and space-time feature parallel extraction is performed, to obtain a multi-modal space-time feature set; based on a dynamic causal discovery algorithm, loop-free causal constraints are performed on the multi-modal space-time feature set, to obtain a modal causal graph; a time delay response kernel matrix is extracted from a lag causal matrix of the modal causal graph, the time delay response kernel matrix is taken as a prior constraint, cross-modal feature fusion is performed on the multi-modal space-time feature set based on a causal attention mechanism, to obtain a fused multi-modal feature; based on a hybrid expert architecture, multi-task risk analysis is performed on the fused multi-modal feature, to obtain a comprehensive risk vector; contribution degree attribution is performed on the comprehensive risk vector, to obtain a contribution heat map, cross verification is performed on the comprehensive risk vector in combination with the modal causal graph, and a tea garden risk report is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of tea garden management and risk analysis technology, and in particular to a method and system for predicting tea garden risks based on multimodal data. Background Technology

[0002] In the field of refined tea garden cultivation and risk management, multimodal data perception and intelligent analysis are the core means to achieve early warning of disasters and ensure tea yield and quality. Modern tea gardens can acquire multi-dimensional and heterogeneous time-series data on meteorology, soil, crop physiology, and human intervention by deploying IoT sensors, remote sensing equipment, and manual recording. These data are characterized by diverse data sources, varying spatiotemporal scales, and complex dynamic evolution, which places extremely high demands on the accurate identification of risk factors, early warning, and causal tracing.

[0003] However, existing tea garden risk prediction methods still face significant technical bottlenecks when dealing with the complex multimodal data and volatile agricultural environments described above. First, most methods rely on single-modal or simple fusion statistical models, lacking the ability to deeply model the spatiotemporal correlations and causal logic implicit in multi-source data. This results in redundant risk discrimination features, high noise interference, and weak interpretability of prediction results. Second, traditional models are mostly based on static correlation analysis, failing to effectively incorporate the temporal dynamic attributes of risk events and unable to characterize the combined impact of environmental factors on the speed and probability of risk occurrence, thus limiting the timeliness and accuracy of predictions. Furthermore, existing methods often separate data-driven models from agricultural physical mechanisms, either over-relying on black-box models with a lack of agronomical logic in the decision-making process, or being limited to empirical rules, making it difficult to adapt to the microclimate variations and sudden disaster scenarios in complex mountain tea gardens. This leads to insufficient stability and poor generalization ability of early warning results in practical applications. Summary of the Invention

[0004] (a) Technical problems to be solved To address the shortcomings of existing technologies, this invention provides a tea garden risk prediction method and system based on multimodal data. It has the advantages of multimodal data causal screening, spatiotemporal dynamic feature fusion, and risk-mechanism cross-validation. It solves the problems of high false alarm rate, lagging time series prediction, and weak mechanism explanation caused by the lack of high-reliability feature denoising, accurate causal association modeling, and decision interpretability assurance in tea garden planting scenarios with high data heterogeneity, strong environmental variability, and complex risk coupling.

[0005] (II) Technical Solution To achieve the above objectives, the present invention provides the following technical solution: This invention provides a tea garden risk prediction method based on multimodal data, comprising the following steps: Based on risk correlation constraints, high-dimensional survival analysis and contribution screening and denoising are performed on the pre-acquired tea garden multimodal dataset to obtain a denoised multimodal dataset. The denoised multimodal dataset is subjected to parallel extraction of spatiotemporal features to obtain a multimodal spatiotemporal feature set. Then, the multimodal spatiotemporal feature set is subjected to acyclic causal constraints based on a dynamic causal discovery algorithm to obtain a modal causal graph. The time-delay response kernel matrix is ​​extracted from the lag causal matrix of the modal causal graph. The time-delay response kernel matrix is ​​used as a prior constraint. Based on the causal attention mechanism, cross-modal feature fusion is performed on the multimodal spatiotemporal feature set to obtain fused multimodal features. Multi-task risk analysis is performed on the fused multimodal features based on a hybrid expert architecture to obtain a comprehensive risk vector; The contribution of the comprehensive risk vector is attributed to obtain a contribution heatmap. The comprehensive risk vector is then cross-validated using the modal causal graph, and a tea garden risk report is generated.

[0006] According to a preferred embodiment of the present invention, a high-dimensional survival analysis and contribution-based denoising are performed on a pre-acquired tea garden multimodal dataset based on risk correlation constraints to obtain a denoised multimodal dataset, including: Based on the risk event data in the pre-acquired tea garden multimodal dataset, survival event transformation and covariate integration are performed on the tea garden multimodal data to obtain a survival analysis dataset; The product limit is estimated on the survival analysis dataset to obtain a set of survival curves. The log-rank test and covariate screening are then performed on the survival analysis dataset based on the set of survival curves to obtain the initial survival dataset. Proportional risk modeling is performed on the initial screening survival dataset, and the contribution of the initial screening survival dataset after proportional risk modeling is quantified based on regularized compression to obtain a modal contribution set. The tea garden multimodal dataset is denoised based on the modal contribution set to obtain a denoised multimodal dataset.

[0007] According to another preferred embodiment of the present invention, the spatiotemporal features of the denoised multimodal dataset are extracted in parallel to obtain a multimodal spatiotemporal feature set, including: The denoised multimodal dataset is temporally reconstructed and spatially mapped to obtain a multimodal spatiotemporal dataset; Temporal features are extracted from the multimodal spatiotemporal dataset to obtain a multimodal temporal feature set; Spatial features are extracted from the multimodal spatiotemporal dataset to obtain a multimodal spatial feature set; Based on the multimodal temporal feature set and the multimodal spatial feature set, spatiotemporal joint modeling is performed on each multimodal spatiotemporal data in the multimodal spatiotemporal dataset to obtain the multimodal spatiotemporal feature set.

[0008] According to another preferred embodiment of the present invention, the multimodal spatiotemporal feature set is subjected to acyclic causal constraints based on a dynamic causal discovery algorithm to obtain a modal causal graph, including: Each multimodal spatiotemporal feature in the multimodal spatiotemporal feature set is used as a graph node to obtain a graph node set. Based on preset prior knowledge and the time lag relationship between each multimodal spatiotemporal feature, causal candidate edges are generated for each graph node to obtain a candidate causal graph. Based on the candidate causal graph, a concurrent causal matrix and a lagged causal matrix are constructed respectively, and the concurrent causal matrix is ​​injected with acyclic constraints based on the exponential trace function to obtain the causal constraint expectation function. The causal constraint expectation function is co-optimized based on the augmented Lagrange method to obtain non-zero causal paths, and the concurrent causal matrix and the lagged causal matrix are updated according to the non-zero causal paths. The candidate causal graph is updated using the updated contemporaneous causal matrix and the lagged causal matrix to obtain the modal causal graph.

[0009] According to another preferred embodiment of the present invention, extracting the time-delay response kernel matrix from the hysteresis causality matrix of the modal causality graph includes: All non-zero elements are extracted from the lagged causal matrix of the modal causal graph to obtain the lagged causal intensity set, and causal correlation feature sequence pairs are extracted from the multimodal spatiotemporal feature set based on the lagged causal intensity set; Calculate the set of cross-relationships between the causal association feature sequence pairs, and extract the time offset from the set of cross-relationships; The time-delay response kernel matrix is ​​calculated based on the hysteresis causality matrix, the cross-correlation set, and the time offset.

[0010] According to another preferred embodiment of the present invention, the time-delay response kernel matrix is ​​used as a priori constraint, and cross-modal feature fusion is performed on the multimodal spatiotemporal feature set based on a causal attention mechanism to obtain fused multimodal features, including: The time-delay response kernel matrix is ​​normalized to obtain a standard response kernel matrix, and the multimodal spatiotemporal feature set is temporally shifted and aligned according to the time offset corresponding to the time-delay response kernel matrix to obtain an aligned spatiotemporal feature set. A linear mapping is performed on the aligned spatiotemporal feature set to obtain a multimodal query vector, a multimodal key vector, and a multimodal value vector; A causal attention mask is constructed based on the contemporaneous causal matrix corresponding to the multimodal spatiotemporal feature set. The attention score between the multimodal query vector and the multimodal key vector is calculated using the causal attention mask as a constraint and the standard response kernel matrix as a prior bias. The causal attention weight is also calculated. The multimodal value vector is weighted and summed based on the causal attention weights to obtain the fused multimodal features.

[0011] According to another preferred embodiment of the present invention, multi-task risk analysis is performed on the fused multimodal features based on a hybrid expert architecture to obtain a comprehensive risk vector, including: Based on the gated network, routing decisions are made using the fused multimodal features to obtain an expert activation weight matrix; The fused multimodal features are nonlinearly mapped using the expert sub-models in the pre-trained hybrid expert pool. The results of each nonlinear mapping are then weighted, fused, residual connected, and normalized according to the expert activation weight matrix to obtain the enhanced multimodal features. The enhanced multimodal features are decoded to obtain a set of predicted risk coefficients. Based on the expert activation weight matrix and the preset risk correlation matrix, the set of predicted risk coefficients is nonlinearly integrated across risk events to obtain a comprehensive risk vector.

[0012] According to another preferred embodiment of the present invention, a contribution heatmap is obtained by attributing the comprehensive risk vector to its contribution, including: Based on the hybrid expert architecture corresponding to the comprehensive risk vector, a benchmark risk vector set corresponding to the preset background dataset is extracted, and the mean of the benchmark risk vector set is used as the risk benchmark value vector. Obtain the multimodal spatiotemporal feature set corresponding to the comprehensive risk vector, and calculate the marginal contribution of the multimodal spatiotemporal feature set relative to the risk benchmark vector based on the model parameters of each expert sub-model to obtain the risk contribution set; Based on the modal causal graph and the time-delay response kernel matrix, the risk contribution set is subjected to causal filtering and time-delay weighting to obtain the spatiotemporal attribution weight set. The spatiotemporal attribution weight set is mapped back to the tea garden grid and the corresponding time axis corresponding to the multimodal spatiotemporal feature set to obtain the contribution heatmap.

[0013] According to another preferred embodiment of the present invention, the comprehensive risk vector is cross-validated in conjunction with the modal causal graph, and a tea garden risk report is generated, including: Based on the modal causal graph, the comprehensive risk vector is logically aligned and causal relationship matched to obtain the risk causal matching matrix. The spatial grid features corresponding to high contribution amounts are extracted from the contribution heatmap, and the spatial grid features are mapped to the topological structure corresponding to the modal causal graph to obtain the risk physical evidence chain; Based on the risk causal matching matrix and the risk physical evidence chain, the comprehensive risk vector is cross-validated and risk correction is performed to obtain the corrected risk vector. The modified risk vector is mapped to risk assessment text and then packaged to generate a tea garden risk report.

[0014] To achieve at least one of the above-mentioned objectives, the present invention further provides a tea garden risk prediction system based on multimodal data. The system includes a data denoising module, which performs high-dimensional survival analysis and contribution screening denoising on a pre-acquired tea garden multimodal dataset based on risk correlation constraints to obtain a denoised multimodal dataset. The causal constraint module performs parallel extraction of spatiotemporal features from the denoised multimodal dataset to obtain a multimodal spatiotemporal feature set, and applies acyclic causal constraints to the multimodal spatiotemporal feature set based on a dynamic causal discovery algorithm to obtain a modal causal graph. The feature fusion module extracts the time-delay response kernel matrix from the lag causal matrix of the modal causal graph, uses the time-delay response kernel matrix as a prior constraint, and performs cross-modal feature fusion on the multimodal spatiotemporal feature set based on the causal attention mechanism to obtain fused multimodal features; The risk analysis module performs multi-task risk analysis on the fused multimodal features based on a hybrid expert architecture to obtain a comprehensive risk vector; The cross-validation module performs contribution attribution on the comprehensive risk vector to obtain a contribution heatmap, performs cross-validation on the comprehensive risk vector in conjunction with the modal causal graph, and generates a tea garden risk report.

[0015] (III) Beneficial Effects Compared with existing technologies, the present invention provides a tea garden risk prediction method and system based on multimodal data, which has the following beneficial effects: This tea garden risk prediction method based on multimodal data introduces risk correlation constraints and uses historical risk events as prior guidance to perform high-dimensional survival analysis and contribution screening on tea garden multimodal data. This effectively avoids the problem of ignoring the time dimension and censoring information in traditional feature screening based on statistical correlation or empirical rules. Through survival event transformation and covariate integration, static risk records are transformed into a time-state joint modeling form, enabling the model to characterize the differences in the impact of different environmental factors on the speed and probability of risk occurrence. Combining product limit estimation and log-rank test, non-parametric screening of risk-sensitive environmental factors is achieved, reducing noise interference in the high-dimensional feature space. Through proportional risk modeling and regularized compression, the contribution of covariates is quantified and sparsified, thereby reducing data redundancy and computational complexity while ensuring risk discrimination capability. This provides a high-confidence and highly interpretable input foundation for subsequent causal structure modeling, multimodal fusion, and risk prediction, significantly improving the stability and generalization ability of risk prediction.

[0016] This tea garden risk prediction method based on multimodal data extracts spatiotemporal features in parallel from a denoised multimodal dataset, unifying heterogeneous and scale-inconsistent multi-source observation data into multimodal spatiotemporal features with clear temporal semantics and spatial structure. This effectively reduces the impact of noise interference and data redundancy on subsequent analysis. Through independent modeling and joint fusion of temporal and spatial features, it can not only depict the temporal evolution of environmental factors but also reflect their diffusion, aggregation, and gradient change characteristics in the spatial dimension, thus forming a higher-order feature expression that is more physically meaningful for the tea garden's operational status. By introducing a dynamic causal discovery mechanism to impose acyclic causal constraints on the multimodal spatiotemporal features, the resulting modal causal graph not only has statistical correlation but also satisfies causal consistency and temporal directionality constraints. This provides an interpretable and verifiable structured causal basis for subsequent risk mechanism analysis, key factor tracing, and intervention strategy optimization, significantly improving the reliability and decision-making value of the overall solution.

[0017] This tea garden risk prediction method based on multimodal data achieves explicit transfer of causal structure information to the feature fusion stage by transforming the lagged causal matrix into a time-delayed response kernel matrix. By jointly modeling the time-delayed response kernel using cross-correlation analysis and causal strength, the method finely characterizes the influence strength and time delay between different modalities, avoiding the problem of neglecting temporal causal constraints in traditional attention mechanisms. By using the time-delayed response kernel as a prior modulation term and constructing a causal attention mask in conjunction with the contemporaneous causal matrix, the cross-modal feature fusion process is ensured to occur only between modal pairs with causal rationality, effectively suppressing spurious correlations and noise interference. This results in fused multimodal features possessing temporal consistency, causal interpretability, and cross-modal synergy, providing a foundation for subsequent risk assessment. The model provides more reliable high-order feature representations for estimation or prediction tasks. By introducing a hybrid expert architecture, it effectively solves the problems of multiple risk types, large modal differences, and complex coupling relationships in multimodal risk analysis of tea gardens. Through a gating network, it can adaptively select the most discriminative expert sub-model based on the fused multimodal features, avoiding the problem of insufficient generalization ability of a single model under different risk scenarios, and improving the overall stability and accuracy of prediction. The sparse activation mechanism significantly reduces computational complexity, enabling the model to maintain high expressive power while having good engineering deployability. By introducing a risk event correlation matrix for cross-risk integration, it can characterize the inherent coupling relationship between risks such as pests and diseases, drought, and yield reduction, avoiding decision bias caused by isolated risk assessment. Attached Figure Description

[0018] Figure 1 The diagram shown is a flowchart of a tea garden risk prediction method based on multimodal data according to the present invention. Detailed Implementation

[0019] The following description is intended to disclose the present invention and enable those skilled in the art to implement it. The preferred embodiments described below are merely examples, and other obvious modifications will occur to those skilled in the art. The basic principles of the invention defined in the following description can be applied to other embodiments, modifications, improvements, equivalents, and other technical solutions that do not depart from the spirit and scope of the invention.

[0020] It is understood that the term "a" should be understood as "at least one" or "one or more", that is, in one embodiment, the number of an element can be one, while in another embodiment, the number of the element can be multiple, and the term "a" should not be understood as a limitation on the number.

[0021] Example 1: Please combine Figure 1 This invention discloses a tea garden risk prediction method based on multimodal data, the method comprising the following steps: Based on risk correlation constraints, high-dimensional survival analysis and contribution screening and denoising are performed on the pre-acquired tea garden multimodal dataset to obtain a denoised multimodal dataset.

[0022] Specifically, since the tea garden multimodal dataset may contain miscellaneous noisy data that is irrelevant or has little correlation with risk analysis, in order to improve the efficiency of risk analysis and reduce the amount of subsequent calculations, it is necessary to filter and remove noise based on the contribution of the data. Before performing high-dimensional survival analysis, this invention also includes the process of data temporal alignment, data cleaning, outlier removal, linear interpolation, and data standardization of the tea garden multimodal dataset, thereby ensuring that the tea garden multimodal dataset used for risk prediction is clean and characteristic data with representativeness.

[0023] The pre-acquired multimodal dataset of tea gardens includes: meteorological data (including ambient temperature, ambient humidity, light intensity, air pressure, rainfall and snowfall, and carbon dioxide concentration) of tea garden areas at different times and locations, obtained through ground IoT sensors, meteorological data, remote sensing data, and planting monitoring logs; soil data (including soil moisture content, soil temperature, soil humidity, soil salinity, and soil conductivity); crop status data (including vegetation index, leaf area index, canopy temperature, and chlorophyll index); human intervention data (including human irrigation events, fertilization events, and pesticide spraying events); and data on risk events that have occurred (including the intensity of pest and disease occurrence, insect population density, frost events, yield reduction events, and drought events).

[0024] In detail, based on risk correlation constraints, high-dimensional survival analysis and contribution-based denoising were performed on the pre-acquired tea garden multimodal dataset to obtain a denoised multimodal dataset, including: Based on the risk event data in the pre-acquired tea garden multimodal dataset, survival event transformation and covariate integration are performed on the tea garden multimodal data to obtain a survival analysis dataset; The product limit is estimated on the survival analysis dataset to obtain a set of survival curves. The log-rank test and covariate screening are then performed on the survival analysis dataset based on the set of survival curves to obtain the initial survival dataset. Proportional risk modeling is performed on the initial screening survival dataset, and the contribution of the initial screening survival dataset after proportional risk modeling is quantified based on regularized compression to obtain a modal contribution set. The tea garden multimodal dataset is denoised based on the modal contribution set to obtain a denoised multimodal dataset.

[0025] The aforementioned survival event transformation and covariate integration refer to transforming risk event data in the tea garden multimodal dataset into survival endpoint labels, and aligning and integrating meteorological data, soil data, crop status data, and human intervention data into a matrix of covariates relative to the survival endpoint labels. For example, "whether yield reduction, frost, drought, or pests and diseases occurred" is defined as... (1 for occurrence, 0 for non-occurrence), the time from the start of monitoring to the occurrence of the risk is defined as the survival time. Samples that have not occurred by the end of the observation period are marked as right-censored. The data of each modality after marking are mapped to the same time scale to construct a high-dimensional feature space matrix, where each row represents a sample of a modality within a survival time interval, and each column represents a covariate. This transforms the static risk record into a dynamic time-state binary vector, enabling the model to handle the speed of risk occurrence. Each survival analysis data in the survival analysis dataset corresponds to a row of covariates, i.e., the survival event data of a modality. The product limit estimation refers to calculating the survival probability corresponding to the occurrence time of each risk event, used to non-parametrically describe the probability evolution of tea gardens maintaining a healthy state under different environmental pressures. The corresponding formula is as follows: in, for Survival rate at any given moment This is a time index, representing the time elapsed since the start of monitoring. This refers to the number arranged in chronological order. The timing of each risk event In order to be in The total number of samples that were still in good health and within the monitoring range before that time. In order to be in At any given time, the number of samples that have experienced a risk event is represented. The survival curves in the survival curve set are generated by product limit estimation and are a series of visual curves with time on the horizontal axis and survival rate on the vertical axis. The log-rank test compares whether there are statistical differences between the survival curves corresponding to different survival analysis data. The log-rank of each survival curve can be calculated using the hypothesis testing operator of the chi-square distribution to preliminarily determine which environmental modalities are sensitive triggers for risk, thus realizing the log-rank test. Covariate screening refers to removing covariates that do not have a significant impact on risk survival rate from the dataset based on the results of the log-rank test.

[0026] Specifically, proportional hazards modeling refers to establishing a proportional hazards function to measure the amplification factor of each covariate on a risky event, while regularization compression refers to introducing a regularization penalty term to implement a sparsity selection operator, resulting in an optimization objective. The optimization objective function used for contribution measurement is as follows: in, For regression coefficients, Indicates the first The occurrence of risk events was observed in a sample of samples. For the index of the covariate, It is the number of covariates. For the first The regression coefficients corresponding to each covariate, when When the corresponding covariate is the risk factor, the risk rate is directly proportional to the increase of the covariate. When the corresponding covariate is the protection factor, the risk rate is inversely proportional to the increase of the covariate. When this occurs, it means that the corresponding covariate is unrelated to risk. It is the first The sampling sample at the ... Observed values ​​for each covariate, It is a risk set, referring to The set of all samples that are in a healthy state (no risk event has occurred) and have not been censored at any given time. The index of the sampled samples within the risk set. It is the first in the risk set The sampling sample at the ... Observed values ​​for each covariate, As a penalty factor, Let be the mixing coefficient, when When using the L1 norm for Lasso regression, the resulting regression coefficients are sparse. When using the L2 norm for Ridge regression, the resulting regression coefficients will be smaller. By quantifying the contribution, the regression coefficients between each risk event and the covariate can be calculated, and the regression coefficients can be aggregated as modal contribution sets for the corresponding modal covariates. Noise removal refers to removing data corresponding to covariates that are not related to the risk events from the Tea Garden multimodal dataset, thereby obtaining a denoised multimodal dataset.

[0027] By introducing risk correlation constraints and using historical risk events as prior guidance, high-dimensional survival analysis and contribution screening are performed on tea garden multimodal data. This effectively avoids the problem of ignoring the time dimension and censoring information in traditional feature screening based on statistical correlation or empirical rules. Through survival event transformation and covariate integration, static risk records are transformed into a time-state joint modeling form, enabling the model to characterize the differences in the impact of different environmental factors on the speed and probability of risk occurrence. Combining product limit estimation and log-rank test, non-parametric screening of risk-sensitive environmental factors is achieved, reducing noise interference in the high-dimensional feature space. Through proportional risk modeling and regularized compression, the contribution of covariates is quantified and sparsified, thereby reducing data redundancy and computational complexity while ensuring risk discrimination capability. This provides a high-confidence and highly interpretable input foundation for subsequent causal structure modeling, multimodal fusion, and risk prediction, significantly improving the stability and generalization ability of risk prediction.

[0028] The denoised multimodal dataset is subjected to parallel extraction of spatiotemporal features to obtain a multimodal spatiotemporal feature set. Then, the multimodal spatiotemporal feature set is subjected to acyclic causal constraints based on a dynamic causal discovery algorithm to obtain a modal causal graph.

[0029] In detail, the denoised multimodal dataset undergoes parallel spatiotemporal feature extraction to obtain a multimodal spatiotemporal feature set, including: The denoised multimodal dataset is temporally reconstructed and spatially mapped to obtain a multimodal spatiotemporal dataset; Temporal features are extracted from the multimodal spatiotemporal dataset to obtain a multimodal temporal feature set; Spatial features are extracted from the multimodal spatiotemporal dataset to obtain a multimodal spatial feature set; Based on the multimodal temporal feature set and the multimodal spatial feature set, spatiotemporal joint modeling is performed on each multimodal spatiotemporal data in the multimodal spatiotemporal dataset to obtain the multimodal spatiotemporal feature set.

[0030] Temporal reconstruction refers to time alignment, windowing, and sequence reorganization of denoised multimodal data according to a unified time benchmark, so that each modality of data forms a temporal sample with a consistent time step and a fixed time span. Timestamp alignment can be used to unify the various denoised multimodal data to a standard step size, such as 1 hour. Spatial mapping refers to remapping the acquisition location, regional grid, or remote sensing image meta-index in the denoised multimodal data according to a unified spatial benchmark to obtain a unified spatial structure representation. Geographic Information System (GIS) coordinate system transformation can be used for spatial mapping. Temporal feature extraction refers to selecting a suitable temporal neural network model to extract temporal features based on the data type of each multimodal spatiotemporal data. For example, for continuous soil data, the soil data is sliced, and long-term environmental fluctuations are divided into discrete semantic blocks. The Informer model is used to extract features such as change trends, fluctuation amplitudes, mutation behaviors, and cumulative effects for each semantic block.

[0031] Specifically, spatial feature extraction refers to modeling the environmental similarity, gradient changes, and spatial dependencies between different spatial locations of various multimodal spatiotemporal data, extracting a subset of spatial features that reflect the propagation, aggregation, or differentiation characteristics of risks in the spatial dimension. For example, when dividing a tea garden area into tea garden grids, when extracting spatial features for soil temperature, not only is the soil temperature of the target grid considered, but the temperatures of adjacent grids are also compared, and spatial operators are used to calculate the temperature gradient and heat diffusion rate. Spatiotemporal joint modeling refers to jointly modeling the interaction relationship between multimodal temporal features and multimodal spatial features belonging to the same multimodal spatiotemporal data, constructing spatiotemporal coupling features that reflect the intensity of the spatial effects of environmental factors under specific time conditions. For example, when abstracting a tea garden area into tea garden grids, multimodal temporal features and multimodal spatial features within each grid are combined, associated, and spliced ​​according to the grid ID corresponding to the spatial features and the timestamp corresponding to the temporal features, thereby obtaining the corresponding multimodal spatiotemporal features, and all multimodal spatiotemporal features are aggregated into a multimodal spatiotemporal feature set.

[0032] In detail, based on the dynamic causal discovery algorithm, acyclic causal constraints are applied to the multimodal spatiotemporal feature set to obtain a modal causal graph, including: Each multimodal spatiotemporal feature in the multimodal spatiotemporal feature set is used as a graph node to obtain a graph node set. Based on preset prior knowledge and the time lag relationship between each multimodal spatiotemporal feature, causal candidate edges are generated for each graph node to obtain a candidate causal graph. Based on the candidate causal graph, a concurrent causal matrix and a lagged causal matrix are constructed respectively, and the concurrent causal matrix is ​​injected with acyclic constraints based on the exponential trace function to obtain the causal constraint expectation function. The causal constraint expectation function is co-optimized based on the augmented Lagrange method to obtain non-zero causal paths, and the concurrent causal matrix and the lagged causal matrix are updated according to the non-zero causal paths. The candidate causal graph is updated using the updated contemporaneous causal matrix and the lagged causal matrix to obtain the modal causal graph.

[0033] The prior knowledge refers to the influence and causal relationships between various modal data obtained in advance. The time lag relationship refers to the lag relationship between various modal data during the period of their influence. For example, meteorological data influences soil data, and this influence has a lag. Therefore, a causal candidate edge pointing from meteorological data to soil data can be generated between the graph nodes corresponding to meteorological data and those corresponding to soil data. Prior knowledge can force the weight of certain impossible causal candidate edges (e.g., future pointing to the past) to always be 0. Constructing the contemporaneous causal matrix and the lag causal matrix refers to arranging the multimodal spatiotemporal feature set according to the contemporaneous time and lag time in a staggered manner to obtain a spatiotemporal feature pair set. All-zero matrix , Represents multimodal spatiotemporal characteristics at the same time. Graph nodes and modal spatiotemporal features The causal strength between graph nodes; constructing a model based on spatiotemporal features for the corresponding features. matrix , Represents the multimodal spatiotemporal characteristics at the lag time. The graph nodes and the modal spatiotemporal features at the current moment The influence strength between graph nodes is represented by the contemporaneous causality matrix, which characterizes the direct causal relationship between different modalities within the same time scale, and the lag causality matrix, which characterizes the temporal causal relationship between modalities at different time scales. The formula for the corresponding causal constraint expectation function is as follows: in, This is a causal matrix for the same period. The diagonal elements are 0. It is a lagged causal matrix. It is the number of samples in the multimodal spatiotemporal feature set. yes The multimodal spatiotemporal feature set corresponding to each time step. yes The multimodal spatiotemporal feature set corresponding to each time step. It is the Frobenius norm. It is a causal matrix of the same period The penalty term coefficient, It is a lagged causal matrix The penalty term coefficient, It is the L1 norm symbol. It is for the causal matrix of the same period Injected acyclic constraints, The trace function symbol, This is the Hadamard product symbol (element-by-element product). This ensures that all elements are squared, thus allowing the trace function to determine the penalty loop. It is the dimension of the covariate; The formula for collaborative optimization is as follows: in, To augment the function of the Lagrange method, For Lagrange multipliers, For penalty parameters, For the loss function; by fixing At Solving using quasi-Newton methods (such as L-BFGS) makes... smallest and and according to The difference from 0 increases with each iteration. And update Finally, the updated version was obtained. and In the process of collaborative optimization, the matrix and Many elements in the matrix will become 0, and the remaining non-zero elements constitute the non-zero causal paths between variables, indicating statistically significant causal relationships in the data. The candidate causal graph is updated using the updated contemporaneous causal matrix and lagged causal matrix to obtain the modal causal graph. For example, when... Then, retain or add a node from the graph. To graph nodes The causal candidate edges, with weights of ,when Then, either retain or add a graph node from the previous time step. Graph nodes up to the current time The causal candidate edges, with weights of .

[0034] By performing parallel extraction of spatiotemporal features from a denoised multimodal dataset, heterogeneous and scale-inconsistent multi-source observation data are uniformly mapped into multimodal spatiotemporal features with clear temporal semantics and spatial structure, effectively reducing the impact of noise interference and data redundancy on subsequent analysis. Through independent modeling and joint fusion of temporal and spatial features, not only can the temporal evolution of environmental factors be characterized, but also their diffusion, aggregation, and gradient change characteristics in the spatial dimension can be reflected, thus forming a higher-order feature expression with more physical meaning for the tea garden's operational status. By introducing a dynamic causal discovery mechanism to impose acyclic causal constraints on the multimodal spatiotemporal features, the resulting modal causal graph not only has statistical correlation but also satisfies causal consistency and temporal directionality constraints. This provides an interpretable and verifiable structured causal basis for subsequent risk mechanism analysis, key factor tracing, and intervention strategy optimization, significantly improving the reliability and decision-making value of the overall solution.

[0035] The time-delay response kernel matrix is ​​extracted from the lag causal matrix of the modal causal graph. The time-delay response kernel matrix is ​​used as a prior constraint. Based on the causal attention mechanism, cross-modal feature fusion is performed on the multimodal spatiotemporal feature set to obtain fused multimodal features.

[0036] Specifically, the time-delay response kernel matrix is ​​extracted from the hysteresis causality matrix of the modal causality graph, including: All non-zero elements are extracted from the lagged causal matrix of the modal causal graph to obtain the lagged causal intensity set, and causal correlation feature sequence pairs are extracted from the multimodal spatiotemporal feature set based on the lagged causal intensity set; Calculate the set of cross-relationships between the causal association feature sequence pairs, and extract the time offset from the set of cross-relationships; The time-delay response kernel matrix is ​​calculated based on the hysteresis causality matrix, the cross-correlation set, and the time offset.

[0037] The non-zero elements correspond to two multimodal spatiotemporal features in the lagged causality matrix that have lagged causal relationships, i.e., the corresponding causal association feature sequence pairs. The cross-correlation set can be calculated using a cross-correlation function. Extracting the time offset means selecting the absolute value of the largest cross-correlation number as the time offset. The corresponding formula is as follows: in, This refers to the multimodal spatiotemporal features in causal association feature sequence pairs. The time delay is Multimodal spatiotemporal characteristics of time The optimal time offset. The time delay step is, i.e. and The time interval between them The preset maximum lag window is set to 72 hours. The time delay step is hour, and The formula for calculating the time-delay response kernel matrix based on the cross-correlation coefficients between the parameters is as follows: in, This refers to the time interval The kernel matrix of time delay response at the th Line 1 The elements of the column, i.e. and In time interval The corresponding time-delay response kernel, In the lagged causality matrix and The corresponding element value, i.e., the lagged causal strength, This is the preset core bandwidth.

[0038] In detail, the time-delay response kernel matrix is ​​used as a priori constraint, and cross-modal feature fusion is performed on the multimodal spatiotemporal feature set based on a causal attention mechanism to obtain fused multimodal features, including: The time-delay response kernel matrix is ​​normalized to obtain a standard response kernel matrix, and the multimodal spatiotemporal feature set is temporally shifted and aligned according to the time offset corresponding to the time-delay response kernel matrix to obtain an aligned spatiotemporal feature set. A linear mapping is performed on the aligned spatiotemporal feature set to obtain a multimodal query vector, a multimodal key vector, and a multimodal value vector; A causal attention mask is constructed based on the contemporaneous causal matrix corresponding to the multimodal spatiotemporal feature set. The attention score between the multimodal query vector and the multimodal key vector is calculated using the causal attention mask as a constraint and the standard response kernel matrix as a prior bias. The causal attention weight is also calculated. The multimodal value vector is weighted and summed based on the causal attention weights to obtain the fused multimodal features.

[0039] The temporal shift alignment refers to resampling and aligning the causal relationship feature sequence pairs in the multimodal spatiotemporal feature set according to a time offset, so that the corresponding two multimodal spatiotemporal features are logically consistent in time phase; for example, for each pair of causal relationship feature sequence pairs... Multimodal spatiotemporal features The sequence is shifted in the future direction. If the step size is inconsistent after translation within a time interval, linear interpolation or aligned sampling is used; linear mapping refers to mapping the aligned spatiotemporal feature set to the attention spaces corresponding to query weights, key weights, and value weights, respectively, thereby generating a multimodal query vector query, a multimodal key vector key, and a multimodal value vector value; constructing a causal attention mask refers to constructing a matrix with the same number of rows and columns as the contemporaneous causal matrix, if Then set the corresponding element in the matrix to ,like Then set the corresponding element in the matrix to The formula for calculating attention score is as follows: in, This refers to the features in the spatiotemporal feature set. With features exist Attention score at any moment It refers to characteristics exist Multimodal query vector at any given time. It refers to characteristics exist The multimodal key vector at time step. It is the transpose symbol. The feature dimension of the multimodal key vector. For learnable causal prior strength coefficients, This refers to the time interval The first standard response kernel matrix Line 1 Column elements, It is the first in the causal attention mask Line 1 The elements of the column; the causal attention weights can be calculated using the softmax operator.

[0040] By transforming the lagged causal matrix into a time-delayed response kernel matrix, the explicit transfer of causal structural information to the feature fusion stage is realized. By jointly modeling the time-delayed response kernel with cross-correlation analysis and causal strength, the influence strength and time delay between different modalities are finely characterized, avoiding the problem of ignoring temporal causal constraints in traditional attention mechanisms. By using the time-delayed response kernel as a prior modulation term and combining it with the contemporaneous causal matrix to construct a causal attention mask, the cross-modal feature fusion process is only carried out between modal pairs with causal rationality, thereby effectively suppressing spurious correlations and noise interference. This allows the fused multimodal features to simultaneously possess temporal consistency, causal interpretability, and cross-modal synergy, providing a more reliable high-order feature representation for subsequent risk assessment or prediction tasks.

[0041] Multi-task risk analysis is performed on the fused multimodal features based on a hybrid expert architecture to obtain a comprehensive risk vector.

[0042] Specifically, based on a hybrid expert architecture, multi-task risk analysis is performed on the fused multimodal features to obtain a comprehensive risk vector, including: Based on the gated network, routing decisions are made using the fused multimodal features to obtain an expert activation weight matrix; The fused multimodal features are nonlinearly mapped using the expert sub-models in the pre-trained hybrid expert pool. The results of each nonlinear mapping are then weighted, fused, residual connected, and normalized according to the expert activation weight matrix to obtain the enhanced multimodal features. The enhanced multimodal features are decoded to obtain a set of predicted risk coefficients. Based on the expert activation weight matrix and the preset risk correlation matrix, the set of predicted risk coefficients is nonlinearly integrated across risk events to obtain a comprehensive risk vector.

[0043] The hybrid expert architecture consists of a gated network and a hybrid expert pool. The hybrid expert pool is a collection of multiple parallel and differentiated feedforward neural networks. Each expert sub-model is a deep neural sub-network. Pre-training refers to using large-scale historical tea garden multimodal data to perform preliminary parameter fitting on the deep neural sub-networks corresponding to each expert sub-model, enabling the identification of risk characteristics under specific modality combinations (e.g., expert 1 is good at identifying "soil data-crop status data" features, and expert 2 is good at identifying "meteorological data-human intervention data" features). The routing decision refers to calculating the gated routing weights of each expert sub-model in the hybrid expert pool for the fusion of multimodal features, calculating the score of each expert through a learnable linear layer, and using the Top-K algorithm to evaluate the results. The scores are filtered (e.g., only the two experts with the highest scores are selected) to achieve sparse activation. Finally, the gated routing weights for each expert sub-model are obtained through Softmax normalization. Residual connection refers to performing residual connection between the weighted fusion result and the original fused multimodal features. Decoding refers to decoding each prediction head in the multi-task prediction head of the hybrid expert pool. Different expert sub-model outputs will enter different prediction heads according to the task type (e.g., pest and disease head, drought head, frost head). The risk correlation matrix is ​​obtained by calculating the correlation between various risk events and performing normalization. Nonlinear integration refers to performing weighted summation and nonlinear transformation on each predicted risk coefficient based on the expert activation weight matrix to obtain a comprehensive risk vector.

[0044] By introducing a hybrid expert architecture, the problems of multiple risk types, large modal differences, and complex coupling relationships in multimodal risk analysis of tea gardens are effectively solved. The gating network can adaptively select the most discriminative expert sub-model based on the fused multimodal features, avoiding the problem of insufficient generalization ability of a single model under different risk scenarios and improving the overall stability and accuracy of predictions. The sparse activation mechanism significantly reduces computational complexity, enabling the model to maintain high expressive power while possessing good engineering deployability. By introducing a risk event correlation matrix for cross-risk integration, the inherent coupling relationships between risks such as pests and diseases, drought, and yield reduction can be characterized, avoiding decision-making biases caused by isolated risk assessments.

[0045] The contribution of the comprehensive risk vector is attributed to obtain a contribution heatmap. The comprehensive risk vector is then cross-validated using the modal causal graph, and a tea garden risk report is generated.

[0046] In detail, the contribution of the comprehensive risk vector is attributed to obtain a contribution heatmap, including: Based on the hybrid expert architecture corresponding to the comprehensive risk vector, a benchmark risk vector set corresponding to the preset background dataset is extracted, and the mean of the benchmark risk vector set is used as the risk benchmark value vector. Obtain the multimodal spatiotemporal feature set corresponding to the comprehensive risk vector, and calculate the marginal contribution of the multimodal spatiotemporal feature set relative to the risk benchmark vector based on the model parameters of each expert sub-model to obtain the risk contribution set; Based on the modal causal graph and the time-delay response kernel matrix, the risk contribution set is subjected to causal filtering and time-delay weighting to obtain the spatiotemporal attribution weight set. The spatiotemporal attribution weight set is mapped back to the tea garden grid and the corresponding time axis corresponding to the multimodal spatiotemporal feature set to obtain the contribution heatmap.

[0047] The background dataset is a multimodal dataset of tea gardens that represents the normal situation of tea gardens. The method for extracting the baseline risk vector is consistent with the method from high-dimensional survival analysis to calculating the comprehensive risk vector in the above steps. The risk contribution set can be calculated by Shapley value decomposition, that is, using the KernelExplainer or DeepExplainer module of the SHAP interpreter, with the background dataset as the baseline, to calculate the SHAP value of each feature dimension in the multimodal spatiotemporal feature set for each risk event in the comprehensive risk vector, and obtain the risk contribution set as the risk contribution. Causal filtering refers to traversing the modal causal graph. If there is no direct edge between two corresponding graph nodes in the modal causal graph, the contribution corresponding to the risk contribution set is attenuated. Time-delay weighting refers to weighting the contribution according to the time-delay causal strength if there is a corresponding lag causal strength in the time-delay response kernel matrix.

[0048] Specifically, the comprehensive risk vector is cross-validated using the modal causal graph, and a tea garden risk report is generated, including: Based on the modal causal graph, the comprehensive risk vector is logically aligned and causal relationship matched to obtain the risk causal matching matrix. The spatial grid features corresponding to high contribution amounts are extracted from the contribution heatmap, and the spatial grid features are mapped to the topological structure corresponding to the modal causal graph to obtain the risk physical evidence chain; Based on the risk causal matching matrix and the risk physical evidence chain, the comprehensive risk vector is cross-validated and risk correction is performed to obtain the corrected risk vector. The modified risk vector is mapped to risk assessment text and then packaged to generate a tea garden risk report.

[0049] Logical alignment and causal relationship matching refer to matching each risk event dimension in the comprehensive risk vector with causal nodes in the modal causal graph to determine whether each risk event has a corresponding direct or indirect causal path, thus obtaining a risk causal matching matrix. For example, logical alignment includes establishing a mapping table between risk events and the set of nodes in the causal graph. If the node corresponding to a risk event has an incoming edge or upstream path in the causal graph, the logical alignment is considered successful, and a matching matrix with binary (0, 1) or continuous values ​​is output. Causal relationship matching refers to retrieving whether there is a causal path pointing to the corresponding risk event in the modal causal graph and determining whether the causal directions are consistent. Extracting spatial grid features refers to extracting the Top-K contribution values ​​or values ​​exceeding a threshold from each risk event in the contribution heatmap. The process involves: 1) extracting specific numerical changes of the corresponding modality within a grid (e.g., soil moisture in grid 3 decreased by 30% in the past 12 hours); 2) generating a risk physical evidence chain, which involves converting these specific grids exceeding the threshold into node state activations in the modal causal graph, thus forming a complete physical evidence chain from "spatial anomaly" to "causal logic" and then to "risk prediction"; 3) cross-validation, which involves calculating the consistency score between the risk causal matching matrix and the risk physical evidence chain; 4) risk correction, which involves maintaining or fine-tuning the weight of risk events with high consistency scores and exponentially decaying the weight of risk events with low consistency scores; and 5) mapping to risk assessment text, which involves inputting the corrected risk vector into a preset logical mapping matrix or a lightweight large language model decoder to output a structured risk assessment text.

[0050] By introducing SHAP-based contribution analysis, the impact of multimodal spatiotemporal characteristics on various risk events can be explained in a quantifiable and traceable manner. By combining modal causal graphs with time-delay response verification of contribution results for causal filtering and time-series weighting, spurious contributions caused solely by statistical correlation can be effectively suppressed. By mapping contribution heatmaps to the tea garden spatial grid and constructing a risk physical evidence chain, a logical closed-loop verification between risk prediction results and actual spatial anomalies is achieved. This allows for weight correction of risk events lacking causal or physical evidence support, thereby improving the credibility, interpretability, and decision reliability of tea garden risk reports.

[0051] Example 2: This invention discloses a tea garden risk prediction system based on multimodal data. The system includes a data denoising module, which performs high-dimensional survival analysis and contribution screening denoising on a pre-acquired tea garden multimodal dataset based on risk correlation constraints to obtain a denoised multimodal dataset. The causal constraint module performs parallel extraction of spatiotemporal features from the denoised multimodal dataset to obtain a multimodal spatiotemporal feature set, and applies acyclic causal constraints to the multimodal spatiotemporal feature set based on a dynamic causal discovery algorithm to obtain a modal causal graph. The feature fusion module extracts the time-delay response kernel matrix from the lag causal matrix of the modal causal graph, uses the time-delay response kernel matrix as a prior constraint, and performs cross-modal feature fusion on the multimodal spatiotemporal feature set based on the causal attention mechanism to obtain fused multimodal features; The risk analysis module performs multi-task risk analysis on the fused multimodal features based on a hybrid expert architecture to obtain a comprehensive risk vector; The cross-validation module performs contribution attribution on the comprehensive risk vector to obtain a contribution heatmap, performs cross-validation on the comprehensive risk vector in conjunction with the modal causal graph, and generates a tea garden risk report.

[0052] The processes described above with reference to the flowcharts in the embodiments disclosed in this invention can be implemented as computer software programs. The embodiments disclosed in this invention include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication component, and / or installed from a removable medium. When the computer program is executed by a central processing unit (CPU), it performs the functions defined in the methods of this application. It should be noted that the computer-readable medium described above in this application can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wire segments, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to: wireless segments, wire segments, optical fibers, RF, etc., or any suitable combination thereof.

[0053] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0054] Those skilled in the art should understand that the embodiments of the present invention described above and shown in the accompanying drawings are merely examples and do not limit the present invention. The purpose of the present invention has been fully and effectively achieved. The functions and structural principles of the present invention have been shown and explained in the embodiments. Without departing from the stated principles, the implementation of the present invention may have any variations or modifications.

Claims

1. A method for tea garden risk prediction based on multi-modal data, characterized in that, The method includes: Based on risk correlation constraints, high-dimensional survival analysis and contribution screening and denoising are performed on the pre-acquired tea garden multimodal dataset to obtain a denoised multimodal dataset. The denoised multimodal dataset is subjected to parallel extraction of spatiotemporal features to obtain a multimodal spatiotemporal feature set. Then, the multimodal spatiotemporal feature set is subjected to acyclic causal constraints based on a dynamic causal discovery algorithm to obtain a modal causal graph. The time-delay response kernel matrix is ​​extracted from the lag causal matrix of the modal causal graph. The time-delay response kernel matrix is ​​used as a prior constraint. Based on the causal attention mechanism, cross-modal feature fusion is performed on the multimodal spatiotemporal feature set to obtain fused multimodal features. Multi-task risk analysis is performed on the fused multimodal features based on a hybrid expert architecture to obtain a comprehensive risk vector; The contribution of the comprehensive risk vector is attributed to obtain a contribution heatmap. The comprehensive risk vector is then cross-validated using the modal causal graph, and a tea garden risk report is generated.

2. The method of claim 1, wherein, Based on risk correlation constraints, high-dimensional survival analysis and contribution-based denoising were performed on the pre-acquired tea garden multimodal dataset to obtain a denoised multimodal dataset, including: Based on the risk event data in the pre-acquired tea garden multimodal dataset, survival event transformation and covariate integration are performed on the tea garden multimodal data to obtain a survival analysis dataset; The product limit is estimated on the survival analysis dataset to obtain a set of survival curves. The log-rank test and covariate screening are then performed on the survival analysis dataset based on the set of survival curves to obtain the initial survival dataset. Proportional risk modeling is performed on the initial screening survival dataset, and the contribution of the initial screening survival dataset after proportional risk modeling is quantified based on regularized compression to obtain a modal contribution set. The tea garden multimodal dataset is denoised based on the modal contribution set to obtain a denoised multimodal dataset.

3. The method of claim 1, wherein, Parallel spatiotemporal feature extraction is performed on the denoised multimodal dataset to obtain a multimodal spatiotemporal feature set, including: The denoised multimodal dataset is temporally reconstructed and spatially mapped to obtain a multimodal spatiotemporal dataset; Temporal features are extracted from the multimodal spatiotemporal dataset to obtain a multimodal temporal feature set; Spatial features are extracted from the multimodal spatiotemporal dataset to obtain a multimodal spatial feature set; Based on the multimodal temporal feature set and the multimodal spatial feature set, spatiotemporal joint modeling is performed on each multimodal spatiotemporal data in the multimodal spatiotemporal dataset to obtain the multimodal spatiotemporal feature set.

4. The method of claim 1, wherein, Based on the dynamic causal discovery algorithm, acyclic causal constraints are applied to the multimodal spatiotemporal feature set to obtain a modal causal graph, including: Each multimodal spatiotemporal feature in the multimodal spatiotemporal feature set is used as a graph node to obtain a graph node set. Based on preset prior knowledge and the time lag relationship between each multimodal spatiotemporal feature, causal candidate edges are generated for each graph node to obtain a candidate causal graph. Based on the candidate causal graph, a concurrent causal matrix and a lagged causal matrix are constructed respectively, and the concurrent causal matrix is ​​injected with acyclic constraints based on the exponential trace function to obtain the causal constraint expectation function. The causal constraint expectation function is co-optimized based on the augmented Lagrange method to obtain non-zero causal paths, and the concurrent causal matrix and the lagged causal matrix are updated according to the non-zero causal paths. The candidate causal graph is updated using the updated contemporaneous causal matrix and the lagged causal matrix to obtain the modal causal graph.

5. The tea garden risk prediction method based on multimodal data according to claim 4, characterized in that, Extracting the time-delay response kernel matrix from the hysteresis causality matrix of the modal causality graph includes: All non-zero elements are extracted from the lagged causal matrix of the modal causal graph to obtain the lagged causal intensity set, and causal correlation feature sequence pairs are extracted from the multimodal spatiotemporal feature set based on the lagged causal intensity set; Calculate the set of cross-relationships between the causal association feature sequence pairs, and extract the time offset from the set of cross-relationships; The time-delay response kernel matrix is ​​calculated based on the hysteresis causality matrix, the cross-correlation set, and the time offset.

6. The tea garden risk prediction method based on multimodal data according to claim 5, characterized in that, Using the time-delay response kernel matrix as a priori constraint, cross-modal feature fusion is performed on the multimodal spatiotemporal feature set based on a causal attention mechanism to obtain fused multimodal features, including: The time-delay response kernel matrix is ​​normalized to obtain a standard response kernel matrix, and the multimodal spatiotemporal feature set is temporally shifted and aligned according to the time offset corresponding to the time-delay response kernel matrix to obtain an aligned spatiotemporal feature set. A linear mapping is performed on the aligned spatiotemporal feature set to obtain a multimodal query vector, a multimodal key vector, and a multimodal value vector; A causal attention mask is constructed based on the contemporaneous causal matrix corresponding to the multimodal spatiotemporal feature set. The attention score between the multimodal query vector and the multimodal key vector is calculated using the causal attention mask as a constraint and the standard response kernel matrix as a prior bias. The causal attention weight is also calculated. The multimodal value vector is weighted and summed based on the causal attention weights to obtain the fused multimodal features.

7. The tea garden risk prediction method based on multimodal data according to claim 1, characterized in that, Multi-task risk analysis is performed on the fused multimodal features based on a hybrid expert architecture to obtain a comprehensive risk vector, including: Based on the gated network, routing decisions are made using the fused multimodal features to obtain an expert activation weight matrix; The fused multimodal features are nonlinearly mapped using the expert sub-models in the pre-trained hybrid expert pool. The results of each nonlinear mapping are then weighted, fused, residual connected, and normalized according to the expert activation weight matrix to obtain the enhanced multimodal features. The enhanced multimodal features are decoded to obtain a set of predicted risk coefficients. Based on the expert activation weight matrix and the preset risk correlation matrix, the set of predicted risk coefficients is nonlinearly integrated across risk events to obtain a comprehensive risk vector.

8. The tea garden risk prediction method based on multimodal data according to claim 7, characterized in that, The contribution attribution of the comprehensive risk vector is performed to obtain a contribution heatmap, including: Based on the hybrid expert architecture corresponding to the comprehensive risk vector, a benchmark risk vector set corresponding to the preset background dataset is extracted, and the mean of the benchmark risk vector set is used as the risk benchmark value vector. Obtain the multimodal spatiotemporal feature set corresponding to the comprehensive risk vector, and calculate the marginal contribution of the multimodal spatiotemporal feature set relative to the risk benchmark vector based on the model parameters of each expert sub-model to obtain the risk contribution set; Based on the modal causal graph and the time-delay response kernel matrix, the risk contribution set is subjected to causal filtering and time-delay weighting to obtain the spatiotemporal attribution weight set. The spatiotemporal attribution weight set is mapped back to the tea garden grid and the corresponding time axis corresponding to the multimodal spatiotemporal feature set to obtain the contribution heatmap.

9. A tea garden risk prediction method based on multimodal data according to claim 8, characterized in that, The comprehensive risk vector is cross-validated using the modal causal graph, and a tea garden risk report is generated, including: Based on the modal causal graph, the comprehensive risk vector is logically aligned and causal relationship matched to obtain the risk causal matching matrix. The spatial grid features corresponding to high contribution amounts are extracted from the contribution heatmap, and the spatial grid features are mapped to the topological structure corresponding to the modal causal graph to obtain the risk physical evidence chain; Based on the risk causal matching matrix and the risk physical evidence chain, the comprehensive risk vector is cross-validated and risk correction is performed to obtain the corrected risk vector. The modified risk vector is mapped to risk assessment text and then packaged to generate a tea garden risk report.

10. A tea garden risk prediction system based on multimodal data, characterized in that, The system includes a data denoising module, a causal constraint module, a feature fusion module, a risk analysis module, and a cross-validation module, wherein: The data denoising module performs high-dimensional survival analysis and contribution screening on the pre-acquired tea garden multimodal dataset based on risk correlation constraints to obtain a denoised multimodal dataset. The causal constraint module performs parallel extraction of spatiotemporal features from the denoised multimodal dataset to obtain a multimodal spatiotemporal feature set, and applies acyclic causal constraints to the multimodal spatiotemporal feature set based on a dynamic causal discovery algorithm to obtain a modal causal graph. The feature fusion module extracts the time-delay response kernel matrix from the lag causal matrix of the modal causal graph, uses the time-delay response kernel matrix as a prior constraint, and performs cross-modal feature fusion on the multimodal spatiotemporal feature set based on the causal attention mechanism to obtain fused multimodal features; The risk analysis module performs multi-task risk analysis on the fused multimodal features based on a hybrid expert architecture to obtain a comprehensive risk vector; The cross-validation module performs contribution attribution on the comprehensive risk vector to obtain a contribution heatmap, performs cross-validation on the comprehensive risk vector in conjunction with the modal causal graph, and generates a tea garden risk report.