Multi-source data space-time fusion analysis method and device, equipment and medium

By performing spatiotemporal fusion processing and feature extraction on multi-source data in agricultural insurance, dynamic status profiles are generated and abnormal events are monitored in real time. This solves the problem of the lack of a unified spatiotemporal fusion mechanism in existing technologies, realizes continuous tracking of the status of target objects and anomaly identification, and improves the intelligence and accuracy of risk identification and claims decision-making.

CN121834706APending Publication Date: 2026-04-10PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-07
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing technologies in agricultural insurance lack a unified spatiotemporal fusion and dynamic monitoring mechanism for multi-source heterogeneous data, making it impossible to continuously track the status of target objects and identify anomalies. This results in insufficient accuracy of risk identification and actuarial models, and low objectivity and efficiency in claims decisions.

Method used

By acquiring multi-source data of temporal imagery and descriptive text for the target area, spatiotemporal fusion processing is performed. The visual processing module extracts temporal features related to the images, and the text processing module extracts text features. Consistency comparison is performed to generate an analysis report, construct a dynamic status profile of the target object, and monitor and identify abnormal events in real time for change detection processing.

Benefits of technology

It enables continuous tracking and anomaly identification of target objects, improving the intelligence and precision of agricultural insurance business in underwriting, claims settlement and risk control, and enhancing the accuracy of risk identification and the efficiency of claims decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121834706A_ABST
    Figure CN121834706A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data analysis, can be applied to business scenes such as financial science and technology, and discloses a multi-source data space-time fusion analysis method, device and equipment and a medium, and the method comprises the steps: obtaining multi-source data containing time sequence image data and description text, and carrying out the space-time fusion to generate standard fusion data, extracting corresponding features by using a visual processing module and a text processing module, performing consistency comparison to generate a first analysis report, constructing a dynamic state file based on time sequence features, writing the dynamic state file into the report, monitoring the file to identify an abnormal event, and after the abnormal event is identified, obtaining time sequence image data before and after the abnormal event is identified, and performing change detection to generate a second analysis report. Through multi-source data fusion and cross-modal feature analysis, a dynamic state file is constructed in combination with time sequence features, and change detection is executed, so that automatic tracking and anomaly recognition of the state of the target object are realized, and the intelligent level of agricultural insurance business analysis and risk prevention and control is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data analysis technology, and in particular to a method, apparatus, equipment and medium for spatiotemporal fusion analysis of multi-source data. Background Technology

[0002] In the agricultural insurance sector within the financial industry, the current agricultural insurance data analysis and management system still suffers from significant structural deficiencies. Traditional methods of obtaining agricultural insurance information primarily rely on manual reporting, offline surveys, and periodic statistics. Data sources are fragmented and lack unified standards, making it difficult to update farmer-reported information, crop types, and insured areas in a timely manner. Due to this information lag, underwriting institutions struggle to obtain real-time, objective production status data, thus affecting the accuracy of risk identification and actuarial models. This management approach, primarily based on static reports, cannot support dynamic risk monitoring and refined pricing, thus hindering financial institutions' ability to accurately underwrite agricultural insurance.

[0003] During the claims process, the risk assessment process after a disaster event relies on manual evidence collection and offline verification, which is labor-intensive and time-consuming, and prone to disputes due to subjective judgment. Existing image recognition methods are mostly limited to single data sources, such as satellite remote sensing or land parcel declaration data, and cannot be combined with multi-dimensional data such as meteorological, geographical, and agricultural monitoring for correlation analysis, making it difficult to accurately identify the time, scope, and impact of a disaster. The lack of systematic processing capabilities for time-series image data also leads to delays and biases in pre- and post-disaster comparative analysis, affecting the objectivity and efficiency of claims decisions.

[0004] Furthermore, financial institutions generally face the problems of data silos and system fragmentation in agricultural insurance management. Existing technologies mostly focus on static identification and single risk assessment, lacking a spatiotemporal fusion mechanism that can continuously monitor dynamic changes in the target area, making it difficult to move risk prevention and control forward. Due to the lack of a unified dynamic status archive system, the underwriting and claims processes cannot form a data loop, and historical comparison records and anomaly detection are difficult to effectively connect. Summary of the Invention

[0005] The main objective of this invention is to provide a method, apparatus, device, and storage medium for spatiotemporal fusion analysis of multi-source data, aiming to solve the technical problem that the existing technology lacks a unified spatiotemporal fusion and dynamic monitoring mechanism for multi-source heterogeneous data, and cannot achieve continuous tracking and anomaly identification of target object status based on joint analysis of images and text.

[0006] To achieve the above objectives, the present invention provides a multi-source data spatiotemporal fusion analysis method, comprising: Acquire multi-source data, including time-series image data and descriptive text, in the target area, and perform spatiotemporal fusion processing on the multi-source data to obtain standard fused data; The standard fusion data is analyzed using the visual processing module to extract temporal features related to the images and obtain the temporal feature vector of the target object. The standard fusion data is analyzed using a text processing module to extract text-related features and obtain the text features of the target object. A consistency comparison is performed between the temporal feature vector of the target object and the textual feature of the target object to generate a first analysis report; A dynamic state profile of the target object is constructed based on the temporal feature vector of the target object, and the first analysis report is written into the dynamic state profile of the target object. Real-time monitoring of the target object's dynamic status profile to identify abnormal events; When an abnormal event is detected, time-series image data before and after the abnormal event is acquired and change detection processing is performed to generate a second analysis report.

[0007] Furthermore, to achieve the above objectives, the present invention provides a multi-source data spatiotemporal fusion analysis device, comprising: The multi-source data fusion module is used to acquire multi-source data including time-series image data and descriptive text in the target area, and to perform spatiotemporal fusion processing on the multi-source data to obtain standard fused data. The visual feature extraction module is used to analyze the standard fusion data using the visual processing module, extract temporal features related to the image, and obtain the temporal feature vector of the target object. The text feature extraction module is used to analyze the standard fused data using the text processing module, extract text-related features, and obtain the text features of the target object. The feature comparison and analysis module is used to perform a consistency comparison between the temporal feature vector of the target object and the textual features of the target object, and generate a first analysis report; The status profile construction module is used to construct a dynamic status profile of the target object based on the temporal feature vector of the target object, and write the first analysis report into the dynamic status profile of the target object. An anomaly monitoring module is used to monitor the dynamic status profile of the target object in real time to identify abnormal events; The change detection module is used to acquire time-series image data before and after the occurrence of an abnormal event and perform change detection processing to generate a second analysis report when an abnormal event is detected.

[0008] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a multi-source data spatiotemporal fusion analysis program stored in the memory and executable on the processor, wherein when the multi-source data spatiotemporal fusion analysis program is executed by the processor, it implements the steps of the multi-source data spatiotemporal fusion analysis method as described above.

[0009] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a multi-source data spatiotemporal fusion analysis program, wherein when the multi-source data spatiotemporal fusion analysis program is executed by a processor, it implements the steps of the multi-source data spatiotemporal fusion analysis method as described above.

[0010] Beneficial Effects: This invention relates to the field of data analysis technology and can be applied to business scenarios such as fintech. It discloses a multi-source data spatiotemporal fusion analysis method, apparatus, device, and medium, including: acquiring multi-source data containing time-series image data and descriptive text in a target area and performing spatiotemporal fusion processing to generate standard fused data; using a visual processing module to extract time-series features related to the images to obtain a time-series feature vector of the target object; using a text processing module to extract text-related features to obtain text features of the target object; performing a consistency comparison between the time-series feature vector of the target object and the text features of the target object to generate a first analysis report; constructing a dynamic state profile of the target object based on the time-series feature vector of the target object and writing it into the first analysis report; monitoring the dynamic state profile of the target object in real time to identify abnormal events; and when an abnormal event is identified, acquiring time-series image data before and after the occurrence of the abnormal event and performing change detection processing to generate a second analysis report. This invention achieves continuous tracking and anomaly identification of the target object's state by performing unified spatiotemporal fusion and cross-modal feature analysis on multi-source heterogeneous data, and constructing dynamic state archives by combining time-series feature vectors, thereby significantly improving the intelligence and precision of agricultural insurance business in underwriting, claims settlement and risk control. Attached Figure Description

[0011] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings: Figure 1 This is a schematic diagram of an application environment for a multi-source data spatiotemporal fusion analysis method according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating an embodiment of the multi-source data spatiotemporal fusion analysis method of the present invention; Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the multi-source data spatiotemporal fusion analysis device of the present invention; Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention; Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation

[0012] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0013] The multi-source data spatiotemporal fusion analysis method provided in this invention can be applied to, for example... Figure 1 In this application environment, the client communicates with the server via a network. The server can obtain multi-source data containing time-series image data and descriptive text in the target area through the client, and perform spatiotemporal fusion processing to generate standard fused data. A visual processing module extracts time-series features related to the images to obtain the time-series feature vector of the target object, and a text processing module extracts text-related features to obtain the text features of the target object. A consistency comparison is performed between the time-series feature vector and the text features of the target object to generate a first analysis report. Based on the time-series feature vector, a dynamic state profile of the target object is constructed and written into the first analysis report. The dynamic state profile of the target object is monitored in real time to identify abnormal events. When an abnormal event is identified, time-series image data before and after the abnormal event is obtained and change detection processing is performed to generate a second analysis report. This invention achieves continuous tracking and anomaly identification of the target object's state by performing unified spatiotemporal fusion and cross-modal feature analysis on multi-source heterogeneous data, combining time-series feature vectors to construct a dynamic state profile and realize real-time monitoring and change detection processing. This significantly improves the intelligence and accuracy of agricultural insurance business in underwriting, claims settlement, and risk control. The client can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will now be described in detail through specific embodiments.

[0014] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the multi-source data spatiotemporal fusion analysis method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.

[0015] like Figure 2 As shown, the spatiotemporal fusion analysis method for multi-source data proposed in this invention includes the following steps: S10, acquire multi-source data including time-series image data and descriptive text in the target area, perform spatiotemporal fusion processing on the multi-source data, and obtain standard fused data; In this embodiment, acquiring multi-source data, including time-series imagery and descriptive text, within the target area is a fundamental step in achieving multi-dimensional information fusion. The target area refers to the analysis object region with clearly defined spatial boundaries, such as an administrative division or a specific crop-growing area. The acquisition process typically relies on the collaborative execution of multiple data acquisition terminals, including remote sensing satellites, drone aerial photography, ground-based IoT sensors, and agricultural reporting systems. Remote sensing satellite data provides continuous monitoring images at a macroscopic scale, covering a large area and possessing stable time-series characteristics; drone aerial photography data supplements high-resolution ground feature details, enhancing spatial resolution; IoT sensors collect real-time environmental parameters such as meteorological and soil conditions, forming auxiliary feature data; descriptive text originates from farmer declarations, disaster reports, or regulatory materials, containing manually recorded structured and unstructured semantic information. These sources differ in time, space, and format; therefore, format standardization and timestamp unification are necessary during the acquisition process to ensure comparability of data from different sources.

[0016] Spatiotemporal fusion of multi-source data aims to address inconsistencies in sampling frequency, geographic coordinate systems, and information types among different data sources. This process typically includes three stages: temporal alignment, spatial alignment, and semantic alignment. Temporal alignment uses timestamp interpolation and frame interpolation algorithms to reconstruct a time-series dataset with a unified time interval from data collected at different times. Spatial alignment uses coordinate mapping and projection transformation algorithms to unify data with different resolutions and coordinate systems under the same geographic coordinate system. Semantic alignment uses feature extraction models to embed and transform the semantic differences between text descriptions and image features, achieving a unified semantic expression. Through multi-layer feature alignment and feature weight balancing mechanisms, multi-dimensional information can be fused within the same data framework. The resulting standard fused data is a composite data structure containing spatial, temporal, and semantic multi-dimensional attributes, directly supporting subsequent visual and text analysis processes.

[0017] In its implementation, spatiotemporal fusion employs a joint architecture of convolutional neural networks and temporal modeling networks to achieve synchronous learning of spatial features and temporal changes. Spatial features are extracted by convolutional layers, while temporal changes are captured through long short-term memory networks or Transformer temporal layers. Temporal correlation of textual information can be achieved through a semantic timestamp matching strategy, binding descriptive text with contemporaneous image frames to form a joint text-image sample set. To ensure fusion accuracy, the system performs consistency checks before each fusion, detecting temporal differences, spatial offsets, and the proportion of missing values ​​in the source data. The fused data is stored using structured identifier indexes, enabling rapid location of data units of specific time, location, and event type during subsequent retrievals.

[0018] In different implementations, data acquisition and spatiotemporal fusion can be achieved through different technical paths. For example, in an implementation combining satellite remote sensing and UAV imagery, a multi-scale image pyramid structure can be used to achieve joint fusion of data at different resolutions, eliminating scale differences through pyramid upsampling and feature overlay. In an implementation incorporating ground sensors, temporal matching of multimodal information can be achieved through temporal interpolation and environmental parameter normalization. In an implementation including text data, a Transformer-based multimodal alignment network can be used to embed textual semantics into the image feature space, thereby achieving semantically consistent spatiotemporal feature fusion. If the system is deployed in a real-time monitoring scenario, a streaming data fusion mechanism can be adopted, enabling sensor and image data to undergo spatiotemporal alignment and standardization immediately after acquisition, improving response speed and timeliness.

[0019] An adaptive weight adjustment mechanism can be introduced into different fusion strategies to dynamically allocate fusion weights based on the confidence level, coverage, and time update frequency of the data source, thereby ensuring that high-confidence data dominates the fusion result. For crop scenes with significant seasonal variations, a time-periodic compensation model can be established to correct the brightness and color of image sequences collected in different seasons, reducing the impact of climate change on feature extraction. For regions with a large amount of missing data, reconstruction methods based on spatial interpolation and generative adversarial networks can be introduced to restore the data continuity of key temporal nodes.

[0020] This embodiment, through the acquisition and spatiotemporal fusion of multi-source data, can construct a unified standard fusion data structure in the time, space, and semantic dimensions, fundamentally eliminating the differences between multiple data sources and improving the accuracy and stability of subsequent visual and text analysis.

[0021] S20, the visual processing module is used to analyze the standard fusion data, extract the temporal features related to the image, and obtain the temporal feature vector of the target object; In this embodiment, analyzing the standard fused data using the visual processing module is a crucial step in achieving image information structuring and temporal feature extraction. After spatiotemporal alignment, the standard fused data possesses a unified temporal index and spatial resolution, thus supporting frame-by-frame analysis by the deep feature extraction model. The visual processing module typically includes an image feature extraction unit, a spatial feature modeling unit, and a temporal correlation analysis unit. The image feature extraction unit is responsible for parsing the image signals in the input data, decoding the multi-source fused image into a standardized image frame sequence; the spatial feature modeling unit extracts local texture and spatial boundary features using convolution operators; and the temporal correlation analysis unit captures the dynamic features of the image over time using a temporal attention mechanism.

[0022] The main task of the image decoding stage is to separate the image components from the fused data structure. Standard fused data typically includes an image matrix, time labels, and geographic coordinate indexes. During reading, the system quickly extracts the image matrix using a data index table and arranges it in chronological order to generate a sequence of feature images. Image equalization and illumination compensation algorithms can be used during the generation of the feature image sequence to eliminate deviations caused by environmental factors such as lighting and weather between images, thereby ensuring the stability of subsequent feature extraction.

[0023] Spatial feature extraction is achieved through a multi-scale convolutional network. Multi-scale convolutional layers extract boundary lines, land cover textures, and target morphology information within different receptive fields, enabling the model to identify plot boundaries, crop canopy structures, and land cover types at different scales. The feature maps output from each convolutional layer are normalized and channel-fused to form spatial feature maps. These spatial feature maps not only preserve local details but also capture contextual information through cross-layer connections, providing an input basis for temporal series analysis.

[0024] The temporal analysis process is implemented using a Transformer encoding structure. The spatial feature map sequence is input into the encoding layer in chronological order. The encoding layer uses a temporal attention mechanism to model the correlation between features in each frame and features in preceding and following frames. Temporal attention weights are used to measure the changing trend of the target object over time, such as changes in vegetation cover, image brightness, or surface humidity. The model incorporates location encoding during processing, allowing temporal information to explicitly participate in feature modeling and ensuring that the output contains temporal dependencies.

[0025] In the temporal feature generation stage, the model aggregates the temporal feature sequences and uses adaptive weighted averaging or time window statistical mechanisms to generate a temporal feature vector for the target object. This vector is a high-dimensional representation that includes the image's temporal trajectory, spatial feature evolution trend, and target state stability. This feature vector can serve as input data for subsequent consistency comparison, anomaly detection, and dynamic archive construction.

[0026] In different implementations, the structure of the vision processing module can be adjusted according to the data source and analysis objectives. For example, when the data mainly comes from satellite remote sensing images, a deep convolutional network combined with spectral feature channels can be used, and the sensitivity of specific bands to crop state changes can be enhanced through a spectral attention layer; when the data source includes UAV images, a lightweight residual convolutional network can be used to adapt to the edge computing environment, and multi-frame difference analysis can be performed after image stitching to extract local change areas.

[0027] In another implementation, the temporal modeling part can be replaced with a gated recurrent unit network to model nonlinear changes in temporal images spanning long periods. A self-supervised contrastive learning mechanism can also be introduced into the temporal modeling to enable the model to automatically capture temporal variation patterns under unlabeled conditions. For different scenarios, dynamic convolution or deformable convolution operators can be introduced in the feature extraction stage to adapt to the spatial non-uniformity of images at different resolutions.

[0028] A multi-task learning framework can be introduced during the model training phase to simultaneously optimize spatial feature extraction and temporal change recognition, thereby improving the discriminative power of temporal feature vectors. For high-dimensional temporal features, a dimensionality reduction encoding mechanism can be used to compress high-dimensional features through autoencoders or principal component analysis methods, retaining the most representative temporal change patterns.

[0029] This embodiment, through analysis by the visual processing module, can extract continuous temporal feature vectors from standard fused data, achieving a comprehensive characterization of the spatial changes and temporal evolution of the target object. This process not only improves the structuring level of image information but also enhances the system's ability to identify long-term dynamic changes, providing high-quality temporal feature input for subsequent consistency comparison and anomaly event identification.

[0030] S30, the text processing module is used to analyze the standard fusion data, extract text-related features, and obtain the text features of the target object; In this embodiment, the purpose of using the text processing module to analyze the standard fusion data is to extract semantic features from unstructured text information and form a semantically aligned feature representation with the image data. The text portion of the standard fusion data typically originates from land parcel declaration information, disaster records, agricultural monitoring reports, regulatory documents, or descriptive materials provided by farmers. This data is presented in natural language and covers aspects such as crop type, growth status, disaster situation, and treatment measures. The task of the text processing module is to extract semantic features related to the target object from these descriptions, enabling the text content to have a quantifiable and computable expression in subsequent consistency comparisons and dynamic archive construction.

[0031] The first stage of text analysis is parsing the semantic information of the text. This parsing process includes word segmentation, syntactic analysis, and entity recognition. Through word segmentation and syntactic tree construction, the system can identify time expressions, geographical entities, and agricultural-related action words, such as "lodging," "drought," and "yield decline." To improve the accuracy of semantic recognition, the system introduces a context-aware mechanism based on a pre-trained language model, enabling it to understand implicit event relationships and descriptive logic within complex semantics.

[0032] After text is input into the language model, the system performs semantic parsing to generate preliminary semantic features. The language model can employ a Transformer architecture, whose multi-head attention mechanism can capture long-distance dependencies, thereby understanding the semantic connections between different sentences. For example, the model can determine the causal relationship between "local disaster" and "decline in production." The preliminary semantic features are represented in high-dimensional vector form, reflecting the semantic distribution and structure of the text content.

[0033] To enhance the semantic expressiveness of text features, the system performs knowledge enhancement processing on the semantic parsing results. This knowledge enhancement relies on a pre-defined domain knowledge base, which covers multi-dimensional entity and relational information related to agriculture, risk control, insurance claims, and meteorological disasters. Through knowledge graph matching and concept expansion mechanisms, the system can map fuzzy descriptions in the text into standardized concepts. For example, when the text contains the phrase "flood inundation," the system can automatically associate it with expanded concepts such as "abnormal water depth" and "crop root damage," making the semantic expression more complete.

[0034] After enhancing semantic features, the system performs semantic vector mapping. This process maps semantic features to a semantic vector space through an embedding layer, giving each text segment a numerical feature representation. The semantic vector space is trained to maintain distance relationships between different semantic categories, ensuring that text segments with similar semantics have higher vector similarity. The final generated target object text feature is a high-dimensional semantic vector capable of representing the object attributes, state changes, and event features implicit in the text description, and is aligned with the image temporal feature vector in subsequent stages.

[0035] In different implementations, the text processing module can adopt different model structures and feature extraction methods depending on the data source and scenario requirements. For example, in scenarios where agricultural insurance application texts are the main data source, a BERT-based semantic coding model can be used. This model captures semantic dependencies through a context window and is suitable for long text processing. In scenarios where disaster reports are the core data source, an event extraction model can be used to extract key information such as disaster time, location, and type through event trigger words and argument recognition technology, forming structured semantic features.

[0036] A multilingual adaptation mechanism can also be introduced in the implementation to support cross-regional monitoring systems. The system uses a language mapping layer to uniformly convert texts in different languages ​​into vector representations in a shared semantic space, thereby achieving unified analysis of multilingual texts. Furthermore, when dealing with text data with high semantic noise, a self-supervised denoising training method can be used to enhance the model's ability to focus on effective semantics through contrastive learning.

[0037] In another implementation, a cross-modal alignment mechanism can be introduced to enhance the correlation between text and image data. This mechanism improves the alignment accuracy between text and image features by jointly training a visual encoder and a language encoder, ensuring that both types of features are distributed in the same semantic space. Furthermore, time-labeling information can be added during the semantic vector mapping stage to imbue text features with temporal attributes, enabling time-correspondence analysis with image temporal features.

[0038] In this embodiment, through the analysis of the text processing module, the system can transform unstructured text information into computable semantic feature vectors, achieving a structured expression of text information. This process effectively solves the problem of traditional text records being difficult to integrate with image data, enabling text data to participate in spatiotemporal analysis and consistency comparison, thereby improving the semantic understanding capability and automation level of the overall analysis system.

[0039] S40, perform a consistency comparison between the temporal feature vector of the target object and the textual feature of the target object, and generate a first analysis report; In this embodiment, comparing the temporal feature vector and textual features of the target object is a crucial process for establishing the logical correspondence between image data and textual data. This comparison aims to verify whether the changes in image features and textual descriptions of the same target object match, thereby determining the accuracy of the recorded content and the authenticity of the event description. The temporal feature vector of the target object represents its dynamic trajectory over time, while the textual features represent the object's state and event attributes described in the textual semantics. The core of the consistency comparison is calculating the similarity relationship between multidimensional features to form a quantifiable consistency index.

[0040] Because temporal features and textual features originate from different sources and have inconsistent dimensional structures and distributions, a mapping layer or alignment function is needed to achieve a unified representation. The mapping layer can employ linear transformation matrices, cross-modal embedding mapping, or attention alignment mechanisms to map image features and textual features to the same feature space. In this space, similar temporal variation patterns and semantic expressions will have high feature similarity.

[0041] The system calculates the similarity values ​​between the temporal feature vector and the textual features of the target object across multiple dimensions in a unified feature space. These dimensions include spatial variation patterns, temporal evolution trends, and semantic description consistency. Cosine similarity, Euclidean distance, or a weighted similarity function based on attention weights can be used for calculation. The calculation results form a multidimensional similarity matrix, with each dimension reflecting a specific type of correspondence, such as spatial distribution consistency or temporal trend consistency.

[0042] The multidimensional similarity results are weighted and fused to obtain a comprehensive consistency score. The weights can be determined through training data statistics or adaptive adjustment mechanisms, enabling the model to dynamically balance the contributions of image features and text features according to different scenarios. For example, in disaster monitoring scenarios, the time dimension can have a relatively high weight to highlight the importance of dynamic trends.

[0043] The system compares the overall consistency score with multiple preset thresholds to determine the consistency level. The levels are categorized into four types: highly consistent, basically consistent, some discrepancy, and significantly abnormal. Different levels represent the degree of matching between image data and text records. This level classification not only provides a basis for subsequent report generation but also offers early warning signals for anomaly event identification.

[0044] The system generates a first analysis report based on the consistency level and comparison results. The report includes a difference analysis matrix, comparison conclusions, anomaly level analysis, and handling recommendations. The difference analysis matrix displays type differences, area differences, and time differences in a structured format, facilitating manual or systematic review; the comparison conclusions summarize the degree of consistency between the image and text descriptions; and the handling recommendations provide operational guidance for subsequent dynamic file updates or anomaly monitoring.

[0045] In different implementations, consistency comparison can be achieved through different feature fusion strategies. For example, in scenarios where satellite imagery and disaster information text are the main data sources, a cross-modal contrastive learning model can be used to train the consistency discriminant function using positive and negative sample pairs; in scenarios where drone imagery and farmer reporting text are the main data sources, a dual encoder architecture can be used, where the image encoder and text encoder are trained separately and then aligned by sharing an embedding space.

[0046] In another implementation, an attention fusion module can be introduced, enabling the system to automatically identify image regions highly correlated with keywords in the text, thereby achieving local consistency comparison. This mechanism is implemented through a cross-attention layer, using text semantics as a query vector to perform weighted aggregation of image feature regions, thus obtaining a more refined semantic matching effect.

[0047] Interpretive mechanisms can also be introduced during the report generation phase, allowing the system to include corresponding matching evidence when outputting analysis results. For example, the analysis report could show which time-series image changes correspond to the textual descriptions of "lodging" or "drought," thereby enhancing the interpretability and credibility of the system's output.

[0048] This embodiment achieves the fusion of image and descriptive information at the semantic level by comparing the consistency of temporal feature vectors and text features, forming a quantifiable consistency index system. This process not only improves the automation of data analysis but also significantly reduces the cost of manual comparison and review, enabling the system to quickly identify anomalies and deviations in data records, providing accurate basis for subsequent dynamic status archive construction and abnormal event identification.

[0049] S50, construct a dynamic state profile of the target object based on the temporal feature vector of the target object, and write the first analysis report into the dynamic state profile of the target object; In this embodiment, the core of constructing a dynamic state archive of a target object based on its temporal feature vector lies in mapping the temporal change pattern into structured archive information, thereby enabling continuous state tracking and data accumulation management of the target object. After generation, the target object's temporal feature vector already contains the change trend and state information of the target object at multiple time points. Through structured modeling, these temporal features can be combined with spatial location, attribute classification, and associated report information to form a unified data archive structure.

[0050] The construction of dynamic status archives first requires spatial partitioning and object identification. The system divides the target area into plots using image segmentation algorithms and generates spatial partition units by combining administrative division matching mechanisms. Each plot unit is assigned a unique identifier for subsequent data retrieval and updates. Plot partitioning is based not only on remote sensing image boundaries but also on auxiliary features such as plot boundaries, river channels, and road distribution in geographic information data to improve boundary recognition accuracy.

[0051] After generating object identifiers, the system uses the temporal feature vectors of the target objects to generate growth curve feature vectors for each plot unit. These growth curve feature vectors, through smoothing and interpolation of the feature sequence over time, reflect the dynamic attributes of the target object at different points in time, such as growth status, color index changes, texture density changes, and environmental adaptability. The system ensures feature comparability between different plots through normalization operations, enabling the archives to perform both horizontal comparisons and vertical tracking.

[0052] After generating the growth curve feature vector, the system establishes a mapping relationship between plot identifiers and corresponding feature vectors. The mapping structure is stored in key-value pair format, supporting fast indexing and dynamic updates. Each time new data is input, the system determines the degree of plot change through feature matching. If a significant difference is detected, a new version node is added to the dynamic status archive, and the associated feature vector is updated.

[0053] The process of writing the first analysis report into the dynamic status archive achieves a deep binding between the archive and the analysis results. The system extracts the consistency comparison results from the first analysis report and correlates them with the growth curve feature vectors of the corresponding land parcel units. The correlated results are embedded in the archive structure as extended fields, including information such as comparison time, grade labels, and difference matrix indexes. To ensure the traceability of the time series, the system introduces a timestamp version control mechanism into the archive, recording the time identifier and source of each update operation, thereby realizing versioned storage of the archive and historical status backtracking.

[0054] The dynamic status archive's structural design supports multi-dimensional queries and visualization. The system can utilize the archive interface to display historical status at the plot level, track reports, anomaly annotation, and feature playback, providing fundamental data support for subsequent anomaly monitoring and risk assessment.

[0055] In different implementations, the construction of dynamic status archives can adopt various data structure designs. A relational database structure can be used, recording parcel identifiers, feature vectors, analysis report indexes, and timestamps in a tabular manner to achieve standardized storage; alternatively, a graph database structure can be used, with parcel units as nodes and growth curve feature vectors and analysis reports as associated edges, enabling graph structure queries of multidimensional data.

[0056] In another implementation, dynamic status archives can be managed using a distributed object storage system. Archive files are stored using parcel identifiers as namespaces, and each archive node contains multiple versions of feature files and associated report summaries, suitable for high-concurrency access to large-scale regional data. The system can also achieve fast comparisons through a hash index mechanism, recording only the differences when archives are updated, thereby reducing storage overhead and improving access efficiency.

[0057] Furthermore, an archive visualization module can be incorporated into the implementation to map the time-series data in the dynamic status archive into curves or heat maps for display, intuitively reflecting the growth trend and consistency changes of the land parcels. This visualization process can be combined with time sliding window control to achieve status retrospection and difference analysis for different time intervals.

[0058] This embodiment constructs a dynamic state profile based on the time-series feature vectors of the target object, achieving structured integration of continuous state management and analysis reports for the target object over time. This mechanism not only enables the persistent storage of multi-source features and comparison results but also supports the tracing, comparison, and dynamic analysis of historical data, providing a long-term data foundation for subsequent anomaly identification and risk decision-making, and improving the intelligence and interpretability of the monitoring system.

[0059] S60, Real-time monitoring of the target object's dynamic status profile to identify abnormal events; In this embodiment, real-time monitoring continuously reads the updated data of the target object's dynamic status archive, and performs time positioning and version comparison between new records and existing versions. Time positioning is based on the timestamp version control information in the archive, establishing a monitoring window covering several recent periods to ensure that each incremental update is included in the comparison sequence under the same time benchmark. Version comparison first verifies data integrity and time consistency, eliminating missing segments, duplicate writes, and out-of-bounds time records, generating a comparable continuous sequence.

[0060] The latest fragment corresponding to the growth curve feature vector is extracted from the updated data and concatenated with historical fragments of the same target object in the archive under the same time alignment rules to form a temporal buffer for detection. The temporal buffer adopts a fixed-length or adaptive-length strategy; the fixed length ensures detection stability, while the adaptive length automatically expands or contracts when data density changes or rhythmic abrupt changes to maintain the reliability of statistics. Within the buffer, multi-granularity statistics are calculated for the growth curve feature vector, including moving mean, moving quantile, rate of change, second difference, shape similarity, and inflection point density, to capture stationary shifts, short-term abrupt changes, and structural morphological changes.

[0061] A normal range threshold is constructed and dynamically adjusted. The normal range threshold is derived from the seasonal baseline of the target object's historical state, the population distribution of neighboring target objects, and similar controls of the same crop or management strategy; these three sources form individual baselines, spatial controls, and similar controls, respectively. The dynamic threshold adjustment employs a weighted fusion and confidence interval expansion mechanism. When environmental conditions or operational rhythms drift significantly, the tolerance interval is increased to reduce false alarms; when multiple sources consistently point to anomalous enhancement, the tolerance interval is tightened to improve detection sensitivity. The threshold system is matched one-to-one with the statistics within the time-series buffer to generate index-level deviation.

[0062] Deviation is scored for anomalies, and alarm trigger conditions are determined. The anomaly score considers three dimensions: single-indicator deviation, inter-indicator consistency, and duration. When a single deviation is high but short-lived, a suppression strategy is employed; when multiple indicators consistently deviate within multiple consecutive monitoring windows, the score is increased. Alarm trigger conditions include threshold exceeding rules, morphological rules, and consistency rules: threshold exceeding rules detect amplitude-based anomalies; morphological rules identify non-linear changes (such as sharp drops after plateaus); and consistency rules verify whether the causal structure between indicators has been disrupted. Hysteresis zones and re-trigger cooling-off periods reduce jitter and repetitive alarms.

[0063] A risk knowledge graph is introduced to perform semantic verification and attribution assistance for candidate anomalies. The risk knowledge graph uses target objects, land parcel units, time intervals, environmental elements, and event types as nodes, linking them to historical anomalies, key points from the first analysis report, and external information. The anomaly scores, deviation indicator sets, and time-space labels of candidate anomalies are mapped to the graph, retrieving possible causal paths and their matching degree with known patterns. When a known pattern supports the current deviation, the confidence level is increased; when the graph indicates conflicting evidence, a pending label is triggered, and the anomaly is entered into a delayed judgment queue.

[0064] Multi-scale consistency checks are performed at both the object and spatial levels. At the object level, the continuity of changes to the same target object is compared between adjacent versions to identify whether abrupt changes have a gradual precursor. At the spatial level, a group consistency check is performed on neighboring plot units. When there is a wide-area synchronous deviation and environmental factors are consistent, it tends to indicate an external shock event; when there is an isolated deviation and the neighborhood is normal, it tends to indicate an individual anomaly or data quality problem. The results of the two-level consistency checks are written back to the current judgment as correction items for anomaly scoring.

[0065] After identifying anomalies, anomaly records are generated and linked back to the target object's dynamic status profile. The anomaly record includes the event type, triggering indicator, impact level, estimated duration, spatial coverage estimate, and decision confidence level. It also includes a buffer fragment hash and a threshold version identifier for traceability. Record writing employs append-only versioning to maintain full replayability; simultaneously, a lightweight summary is generated for online monitoring to support fast querying and large-scale retrieval.

[0066] The monitoring strategy and parameters are maintained continuously during operation. The threshold for anomaly scoring, monitoring window length, and weighted fusion parameters are updated online through learning. Update signals are derived from subsequent second analysis report conclusions, historical comparison records, and overall false positive rate metrics across objects. Parameter updates employ small incremental increments and boundary constraints to avoid significant strategy drift caused by short-term fluctuations. All parameter changes generate metadata and are bound to version control information to ensure auditability and reproducibility.

[0067] This embodiment achieves real-time identification and reliable output of abnormal events by continuously reading the dynamic status file of the target object, verifying temporal consistency, constructing a time-series buffer, fusing and dynamically correcting multi-source thresholds within the normal range, multi-dimensionally determining anomaly scoring and alarm conditions, and performing semantic verification and spatial consistency review of the knowledge graph. This mechanism enhances temporal continuity and spatial correlation without increasing the burden of field evidence collection, reduces false positives and false negatives, supports playback and auditing, and can gradually optimize monitoring parameters based on operational data, thereby improving the timeliness, stability, and interpretability of monitoring in financial risk control scenarios.

[0068] S70, when an abnormal event is detected, the time-series image data before and after the abnormal event is acquired and change detection processing is performed to generate a second analysis report.

[0069] In this embodiment, when an abnormal event is detected, the system locates the time period of the abnormal event based on the timestamp record of the target object's dynamic status archive and retrieves the time-series image data before and after that time period. The retrieval process relies on the unique identifier of the land parcel and version control information recorded in the archive to obtain the corresponding image data group from the image storage index. To ensure the accuracy of the comparison, the system performs time synchronization and spatial registration operations to unify the pre-disaster and post-disaster images to the same coordinate reference and spatial resolution, forming a paired image sequence. Time synchronization adopts a linear interpolation and frame alignment mechanism. When there are time intervals or inconsistent sampling frequencies in the image data, interpolation is used to generate transition images to ensure temporal continuity. Spatial registration is based on feature point matching and affine transformation models. By identifying ground boundary, road node, or highly reflective targets as anchor points, the geometric transformation matrix is ​​calculated to eliminate spatial displacement and rotation errors.

[0070] After image registration is completed, the system enters the change detection phase. Change detection processing uses first and second time-series image data as input to calculate change indicators reflecting spatial and spectral variations. Commonly used indicators include differences in normalized vegetation index, changes in spectral correlation coefficient, changes in texture entropy, and changes in reflectance gradient. The system selects different detection models based on image type and monitoring target: for crop growth monitoring, a change detection model based on spectral indices can be used; for surface structure change detection, a detection model based on principal component difference or convolutional feature difference can be used.

[0071] The system performs normalization and noise suppression on the change detection indicators. Normalization eliminates biases introduced by differences in image acquisition conditions at different times, while noise suppression reduces local noise interference through bilateral filtering or multi-scale median filtering. Subsequently, the system compares the normalized change detection indicators with preset thresholds to determine the significance label of each pixel's change. The thresholds can be calculated based on historical sample statistics or adaptive distributions; the adaptive thresholds are automatically adjusted across different land cover types to improve detection accuracy.

[0072] After determining the significance of the change, the system performs aggregation analysis on the changed pixels within the spatial neighborhood to generate a candidate set of changed regions. The aggregation analysis, based on connected component labeling algorithms and spatial morphological operations, removes isolated noise points and extracts continuously changing regions. For each changed region, the system calculates its area, shape complexity, and average change magnitude, forming a spatial change feature vector. The system inputs these features into the impact level determination module, which determines the impact level by calculating similarity with historical change patterns or based on cluster center distance. The classification of impact level considers not only the change magnitude but also the duration and spatial coverage of the changed region to reflect the overall scope of the event's impact.

[0073] Once the impact severity level is generated, the system constructs preliminary impact analysis results based on the change indicators and level results. These results include the location of the change, the time range, the type of change, and the confidence level. The confidence level is calculated by considering the consistency of multidimensional features and the level of detection noise. If the confidence level is lower than a preset threshold, the system triggers a review process, confirming the stability of the results through resampling or model re-detection. Finally, based on the preliminary impact analysis results and the review results, the system generates a second analysis report. The report includes a summary of the change detection indicators, the impact severity level, a spatial distribution map, and processing recommendations. The report generation uses a structured template to support subsequent database archiving and strategy invocation.

[0074] This embodiment achieves quantitative analysis and traceable recording of changes in the state of a target object by automatically extracting time-series image data and performing change detection processing after an anomaly occurs. This process combines temporal continuity with spatial consistency, enabling accurate identification of the scope and intensity of the impact of anomalies on the target object, avoiding the uncertainty of manual judgment. Through dynamic threshold and confidence level verification mechanisms, the system effectively suppresses misjudgments caused by environmental noise, improving the stability and reliability of change detection. The resulting second analysis report not only provides objective evidence for risk decision-making and subsequent handling but also provides precise data support for version updates and historical comparisons of the dynamic state archive, thereby realizing an automated, refined, and verifiable anomaly response mechanism.

[0075] This invention relates to the field of data analysis technology and can be applied to business scenarios such as fintech. It discloses a multi-source data spatiotemporal fusion analysis method, apparatus, device, and medium, comprising: acquiring multi-source data containing time-series image data and descriptive text in a target area and performing spatiotemporal fusion processing to generate standard fused data; using a visual processing module to extract time-series features related to the images to obtain a time-series feature vector of the target object; using a text processing module to extract text-related features to obtain text features of the target object; performing a consistency comparison between the time-series feature vector and the text features of the target object to generate a first analysis report; constructing a dynamic state profile of the target object based on the time-series feature vector and writing it into the first analysis report; performing real-time monitoring of the dynamic state profile of the target object to identify abnormal events; and when an abnormal event is identified, acquiring time-series image data before and after the occurrence of the abnormal event and performing change detection processing to generate a second analysis report. This invention achieves continuous tracking and anomaly identification of the target object's state by performing unified spatiotemporal fusion and cross-modal feature analysis on multi-source heterogeneous data, and constructing dynamic state archives by combining time-series feature vectors, thereby significantly improving the intelligence and precision of agricultural insurance business in underwriting, claims settlement and risk control.

[0076] In one embodiment, step S10 above includes: S101, Collect time-series image data of the target area, including satellite remote sensing image data and UAV aerial image data; S102, Collect descriptive text including land parcel declaration information and disaster reports in the target area; S103, Collect IoT sensor data including meteorological data and soil data in the target area; S104, Collect geographic information data including administrative divisions and land parcel boundaries in the target area; S105, perform time alignment processing on the time-series image data, the descriptive text, the IoT sensor data and the geographic information data to obtain time-aligned data; S106, Perform spatial alignment processing on the time-aligned data to obtain spatially aligned data; S107, Extract the spatiotemporal features of the spatial alignment data to obtain spatiotemporal feature data; S108, Generate standard fusion data based on the spatiotemporal feature data.

[0077] In this embodiment, the process of acquiring multi-source data, including time-series imagery and descriptive text, in the target area begins with a unified data acquisition list and establishes an acquisition strategy around three constraints: timestamps, spatial references, and data quality. Time-series imagery comes from satellite remote sensing imagery and UAV aerial imagery. The former uses a preset orbital transit plan and cloud cover filtering rules to determine available image windows, while the latter uses flight path planning and overlap control to ensure continuous coverage and consistency with ground resolution. Before being stored, both types of imagery undergo file verification, image header information parsing, and radiometric calibration to form image metadata records containing sensor model, resolution, observation time, solar altitude angle, and viewing angle information, used for subsequent registration and illumination geometry correction. Descriptive text sources include land parcel declaration information and disaster reports. During the acquisition phase, both types of text are marked with source tags, time stamps, and spatial association tags. A unified denoising and encoding strategy is used to unify texts of different formats into structured text entries with reliability tags and reference paths to ensure traceability and auditability. The IoT sensor data covers meteorological and soil data. The acquisition process requires complete equipment identification, sensor range, sampling frequency, missing measurement markers, and calibration coefficients. Before being stored, abnormal peaks, long zero-value intervals, and drift phenomena are removed and corrected, and the correction coefficients are retained to support verification. Geographic information data includes administrative divisions and land parcel boundaries. Before entering the data warehouse, coordinate system self-checks, topological consistency checks, and boundary closure repairs are performed. A hierarchical mapping relationship from administrative divisions to land parcels is established to ensure subsequent spatial aggregation and authorization isolation.

[0078] The time alignment process aims to construct a unified timeline. First, it converts the timestamps of all data to the same standard time and reliably fills in missing timestamps. Image time alignment employs a windowed resampling strategy, mapping satellite remote sensing imagery and UAV aerial imagery to a unified time scale. When unequal intervals exist between observations, representative time images are generated by calculating interpolation weights and the confidence level of adjacent time images, and the sampling start and end ranges are recorded to avoid information confusion. Descriptive text time alignment utilizes a selective fusion of three time indicators: document generation time, event occurrence time, and reporting time. When these three are inconsistent, the event occurrence time takes precedence, and the deviation is recorded using a confidence factor. IoT sensor data and meteorological data undergo alignment window evaluation in the time dimension, generating statistics with a unified time granularity using minute-level or hour-level sliding windows and labeling observation coverage and missing rate. Geographic information data establishes a valid time interval using version time and effective zoning time to ensure zoning consistency during historical backtracking. After completing the unified timeline, time-aligned data is generated, and metadata records of the alignment strategy, window width, and confidence level are retained.

[0079] Spatial alignment processing aims to unify the spatial benchmark. First, coordinate system one is executed, projecting satellite remote sensing imagery and UAV aerial imagery onto a selected projection. Then, fine-grained registration is performed under corresponding feature constraints, using road intersections, waterway inflection points, and permanent corner points as control points. Geometric transformations are estimated using polynomial or affine models, and registration residual maps are output to quantify alignment accuracy. Pyramid resampling is used to establish a multi-scale image stack for images of different resolutions, ensuring stable calculations at both the plot and pixel scales. Spatial connections are performed between IoT sensor locations and plot boundaries, using nearest neighbor or inverse distance weights to perform spatial interpolation from stations to plots, recording interpolation kernel parameters and the number of valid stations. At the topological level, overlap detection, gap repair, and multi-polygon merging are performed between administrative divisions and plot boundaries to form a topologically consistent set of regions. This yields spatial alignment data, and spatial error assessment values ​​are output for subsequent quality control.

[0080] Spatiotemporal feature extraction constructs an index matrix around spatially aligned data to form spatiotemporal feature data. On the image side, spectral, texture, and deformation features are extracted. Spectral features include the Normalized Difference Vegetation Index (NDVI), soil and water indices, and red-edge features. Texture features are calculated using the gray-level co-occurrence matrix and multi-scale directional gradients to determine local structure. Deformation features are calculated by differentiating and trend slopes on multi-temporal images. On the IoT sensor side, window statistics, cumulative values, and extreme values ​​of temperature and precipitation are extracted. On the soil side, water content fluctuation amplitude and rewetting lag indicators are extracted. On the geographic information side, plot area, shape index, boundary complexity, and adjacency relationships are calculated, and neighborhood aggregation features are generated. On the text side, event labels, time windows, and spatial anchors are obtained from parsing previously structured text entries and projected as a sparse event sequence on the time axis. Then, the event count within the time window, event intensity weights, and event type distribution are used to form the text-side temporal features. All features are subjected to a unified feature alignment and scale normalization process. Intra-batch standardization or quantile mapping is used to control scale differences. Missing test locations are masked and filled in with a reconstructible imputation strategy. At the same time, feature quality scores are generated to record confidence intervals and source coverage, forming spatiotemporal feature data.

[0081] The generation of standard fusion data is constrained by a unified schema definition and traceable metadata. Multi-source spatiotemporal features are organized into a fusion table oriented towards learning and statistics through a key-value structure and columnar storage. The primary key consists of a land parcel identifier and a timestamp, while the extended keys include administrative level, sensor source, and image batch. Each column corresponds to a single feature and is accompanied by a quality score, source label, and transformation parameter hash to reproduce the feature generation process. During the fusion process, correlation screening and redundancy control are performed, removing highly linearly correlated columns and retaining representative principal components to reduce redundancy. Screening thresholds and a list of removed columns are recorded. Upon completion, standard fusion data is output, while retaining processing logs, parameter lists, version tags, and data checksums, providing a consistent entry point for subsequent visual and text processing modules to directly read the data.

[0082] This embodiment introduces unified metadata and quantifiable error records in data acquisition, temporal alignment, and spatial alignment to avoid implicit biases in cross-source data across time and spatial benchmarks, thereby improving the effective sample density and comparability of downstream analysis. By using a multi-dimensional spatiotemporal feature system and quality scores for parallel output, the subsequent learning process can allocate weights based on confidence levels and reduce noise amplification effects. Through the patterned organization and versioned recording of standardized fusion data, a verifiable and traceable fusion result is formed, enabling visual processing and text processing to work together under the same data entry point, improving feature consistency and processing efficiency. Furthermore, redundancy control and source marking reduce storage overhead and improve query performance and batch processing throughput.

[0083] In one embodiment, step S20 above includes: S201, decode the image feature information from the standard fusion data; S202, reconstruct the image feature information into a feature image sequence; S203, The feature image sequence is processed by the multi-scale convolutional layer in the visual processing module to extract multi-scale spatial features and obtain a spatial feature map; S204, The spatial feature map is input into the Transformer encoder of the vision processing module, and the spatial feature map is processed by the temporal attention mechanism to obtain the temporal feature sequence; S205, Based on the time-series feature sequence, generate the time-series feature vector of the target object.

[0084] In this embodiment, when the vision processing module decodes image feature information from standard fused data, it first locates the image source, imaging time, geometric reference, and radiometric calibration parameters based on the metadata of the fused data. It then reads the spectral channels by image batch and applies radiometric and atmospheric correction coefficients to restore the physical reflectance level. Simultaneously, based on cloud and shadow masks, it generates effective pixel masks and quality markers at the pixel level, constructing image feature information entries containing pixel values, masks, and quality weights. To ensure consistency across sources, simplified illumination normalization and terrain correction are performed based on sensor perspective and solar geometry information, compressing brightness differences between different batches caused by observation conditions to a controllable range. The processing parameter hashes are then written into the entries to maintain traceability.

[0085] When reconstructing image feature information into a feature image sequence, pixels are aggregated according to plot grids or regional grids using the effective observation scale on the time axis as an index, and a tile sequence with fixed resolution and fixed size is established. Missing frames or low-quality frames are processed jointly with quality weights and interpolation weights. Proxy tiles are generated by using image candidates from nearby time points and quality weights, while retaining frame-level missing masks so that subsequent attention mechanisms can skip invalid positions. To consolidate cross-scale representation, a multi-scale isotopic pyramid is generated during the reconstruction stage. Each frame contains paired tiles of the original scale and the downsampled scale, ensuring that spatial details and contextual structure are simultaneously included in subsequent calculations.

[0086] When processing feature image sequences, multi-scale convolutional layers first independently execute combinations of convolutional blocks, normalization, and nonlinear units at each scale to extract edges, textures, and shape patterns. Then, dilated convolutions and residual connections expand the receptive field, reducing long-range dependency gaps within the same scale. Next, lateral connections and top-down fusion are introduced between scales to align high-level semantics with low-level details and merge them step-by-step, forming multi-scale spatial features containing semantic density and geometric details. To improve the preservation of ground object boundaries and slender targets, orientation-sensitive kernels and channel attention are introduced into the convolutional blocks, giving higher responses to fine structures such as water systems, roads, and striped crops. After the above processing, a spatial feature map is generated at each time frame, and a frame-level spatial quality score is output simultaneously for weight scheduling in subsequent temporal modeling.

[0087] When the spatial feature map enters the Transformer encoder, it first constructs a composite representation of spatial and scale location information by overlaying row and column coordinates with multi-scale indexes using positional encoding. Then, it generates an attention mask and attention weights based on the effective pixel mask and quality score to mask cloud or ghosting areas and reduce the interference of low-confidence areas on attention allocation. Within a single frame, the encoder aggregates regional context using local attention or windowed attention. In the sequence dimension, it connects adjacent and cross-seasonal frames using a temporal attention mechanism. It uses significant decay and interval sampling strategies to suppress the accumulation of long-distance noise and enhances change sensitivity through temporal difference key value design, so that dynamic processes such as growth, harvesting, water accumulation, lodging, and bare ground exposure form stable trajectories in the attention map. To stabilize long sequence computation, the encoder introduces gradient checkpoints and key value buffers between layers and uses mask alignment rules to ensure that there is no drift between the supplementary frame and the real frame when querying key value pairing.

[0088] After the temporal feature sequence is generated, statistical and trend aggregation along the time dimension is first performed to extract components such as long-term trends, periodic fluctuations, and abrupt gradients. Then, a spatial consistency score is calculated based on plot boundaries and neighborhood relationships, transforming the consistency of pixels within the same plot and the contrast across plot boundaries into time-weighted factors. Next, attention pooling or gated pooling is used to aggregate the contributions of different time slices along the time axis, forming a temporal representation sensitive to seasonal phases, extreme weather responses, and farming behavior. Finally, under feature normalization and regularization constraints, the temporal feature vector of the target object is output, along with source tags, time coverage intervals, effective frame ratios, and mask statistics, for direct use in subsequent consistency comparisons and dynamic archive construction.

[0089] This embodiment reduces the impact of clouds, imaging angles, and unequal interval observations on representation stability by introducing quality weights, mask propagation, and cross-scale alignment in decoding, reconstruction, convolutional extraction, and temporal modeling. This makes the temporal representation transformed from static frame stacking synchronously sensitive to continuous growth processes and sudden changes. Through the synergy of spatial feature maps and temporal attention mechanisms, it takes into account detail boundaries and cross-seasonal evolution, and converges the differences between remote sensing and aerial photography sources into a unified temporal vector expression. This improves the discriminative power and traceability of subsequent text alignment, consistency comparison, and anomaly identification, and supports robust inference in low-quality observation scenarios with masks and source markers.

[0090] In one embodiment, step S30 above includes: S301, Extract textual semantic information from the standard fused data; S302, The text semantic information is input into the language model of the text processing module, and the language model is used to perform semantic parsing on the text semantic information to obtain preliminary semantic features; S303, Based on a preset domain knowledge base, perform knowledge enhancement processing on the preliminary semantic features to obtain enhanced semantic features; S304, the enhanced semantic features are mapped to the semantic vector space to generate text features of the target object.

[0091] In this embodiment, the text processing module first locates the text source and time tag for standard fused data. When parsing the text semantic information, it aligns text fragments with corresponding spatial objects or business objects using the metakeys of data records, unifies character encoding and word segmentation standards, removes control characters and repeated punctuation, and retains the original paragraph and list structure as clues for subsequent hierarchical levels. For fragments containing mixed languages ​​and industry abbreviations, it establishes lexical reconstruction mapping and alias tables. It extracts specific names, quantitative expressions, time limit expressions, place names, and crop names from source fields such as "land parcel declaration information" and "disaster report" as candidate tags through a combination of rule templates and statistical dictionaries, and writes them into the text semantic information using source confidence and timestamp confidence as two additional channels. To reduce the impact of noise, a mask sequence is constructed using source quality scores and character-level defect markers to mask encoding anomalies or truncated fragments while retaining key information bits.

[0092] When textual semantic information is fed into the language model, the contextual location signal is first fused from intra-segment positional encoding and inter-segment source encoding, embedding geographic unit identifiers, time intervals, and data source types into the same input stream. For long texts, a sliding window overlap strategy is employed, and cross-window cache keys are maintained to ensure that cross-segment citations and cross-sentence references remain continuous within the attention space. For tabular descriptions or semi-structured templates, row and column hints are placed using layout markers, giving the model structural priors when decoding implicit fields. During the encoding phase, the language model captures long-range dependencies and parallel relationships through multi-head attention, projects them onto the intermediate semantic channel through a feedforward network, and outputs preliminary semantic features containing trigger words, predicates, arguments, and their dependencies. Simultaneously, each relationship is assigned a source channel weight and a cross-segment consistency score to facilitate subsequent knowledge alignment.

[0093] When performing knowledge enhancement processing on preliminary semantic features based on a pre-defined domain knowledge base, the process first involves standardizing entities to align industry vocabularies with ontology concepts, merging synonyms, colloquialisms, and coded names into a unified identifier. Then, relation alignment is used to perform subgraph matching between patterns such as "crop-growth period-management behavior," "event-time-location," and "disaster cause-manifestation-affected object" and ontology relation edges. Fragments that fail to match are entered into alias expansion and spelling correction loops, while successfully matched fragments generate entity-relation-attribute triples and are assigned ontology-level path and attribute constraints. Subsequently, knowledge graph embedding is used to supplement the preliminary semantic features with adjacent concepts and inferred attributes. For example, a reasonable growth period range is inferred from crop type and regional climate zone, and possible impact types and duration intervals are inferred from disaster cause words and time periods. A conflict detection strategy is used to identify mutually exclusive statements in the text, and conflict markers are injected back into the semantic channel as suppression coefficients. When there are texts from multiple sources, the same semantic slot is weighted and merged based on source weight and historical consistency score, outputting enhanced semantic features while maintaining complete metadata in three aspects: ontology alignment identifier, inference path summary, and confidence interval.

[0094] When enhancing semantic features and mapping them to the semantic vector space, a concatenated representation composed of entity vectors, relation vectors, attribute vectors, and temporal location vectors is first constructed. Hierarchical encoding is used for concepts with hierarchical attributes to maintain the separability of parent-child relationships. For events with time spans, a combination of start-end vectors and interval center vectors is used to express temporal persistence and suddenness. To make texts from different sources and of different lengths comparable in the same space, intra-domain batch normalization and contrastive learning constraints are introduced. This brings the vectors of the same object closer together in different texts and pulls semantically conflicting objects further apart. A mask alignment strategy is used to ignore the impact of missing slots on alignment loss. Finally, the target object text features are obtained through linear transformation and normalization. Simultaneously, three traceable labels are output: source aggregation weight, covered ontology path, and time coverage interval. This ensures that subsequent consistency comparisons with time-series vectors and dynamic file writing can be directly referenced. Throughout the process, textual semantic information, preliminary semantic features, enhanced semantic features, and target object text features maintain a one-to-one mapping relationship in the data structure. Any merging, error correction, or deduction is solidified through change records and version hashes, facilitating audit playback and differential updates.

[0095] This embodiment introduces source alignment, structural tagging, domain ontology, and contrast constraints at each stage of parsing, encoding, knowledge alignment, and vectorization. This transforms the original text expression into a vector representation that can be directly associated with spatial and temporal objects, weakening the interference of differences in terminology and abbreviations on semantic judgment, and improving the alignment stability and comparability of cross-source and cross-template texts. Through knowledge enhancement, explicit descriptions and implicit backgrounds are linked to a unified ontology, missing slots are filled, and conflicts are marked. This provides verifiable semantic support and a traceability path for subsequent comparisons with temporal vectors, thus balancing separability and interpretability when generating text features of target objects, and providing a high-confidence text-side representation foundation for subsequent consistency comparisons, anomaly detection, and dynamic file management.

[0096] In one embodiment, step S40 above includes: S401, determine the similarity values ​​between the temporal feature vector of the target object and the textual features of the target object in multiple feature dimensions, and obtain a multidimensional similarity result; S402, perform weighted fusion processing on the multidimensional similarity results to obtain a comprehensive consistency score; S403, compare the overall consistency score with multiple preset thresholds to determine the consistency level; S404, Generate a difference analysis matrix containing analysis results of type difference, area difference and time difference based on the consistency level; S405, Based on the difference analysis matrix and the consistency level, generate a first analysis report containing detailed comparison conclusions, anomaly level analysis and handling suggestions; S406, perform correlation analysis between the first analysis report and historical comparison records, and update the benchmark parameters for consistency comparison.

[0097] In this embodiment, for the paired input of the target object's temporal feature vector and the target object's textual features, a dimensional mapping relationship is first established, mapping the image-related temporal semantic slots to the textual semantic slots one-to-one. This includes four main dimensions: object category, spatial coverage, time interval, and contextual attributes. Source weight, coverage, and freshness are used as three supplementary channels written into the same paired record. Scale normalization and time axis resampling are performed on the target object's temporal feature vector to ensure the comparability of amplitudes and rhythms at different sampling frequencies. Slot completion and conflict suppression are performed on the target object's textual features, aggregating multiple source text entries in the same semantic slot into a single vector representation according to source weights, while retaining conflict markers for unaggregated entries for subsequent confidence control.

[0098] The calculation of multidimensional similarity results is completed in a shared alignment space. The object category dimension uses a metric learning projector to map the two-sided vectors to the discriminative subspace, extracts category consistency signals using angular similarity, and uses a confusion control set as a suppression term to reduce misclassification of co-domain neighbors. The spatial coverage dimension reads the period-by-period coverage encoding and geometric boundary embedding from the target object's temporal feature vectors, and reads the area and boundary semantic vectors from the target object's textual features. After spatial index alignment, spatial similarity is evaluated in parallel using a region overlap kernel and a boundary consistency kernel. Simultaneously, adjacency constraints of plot topology are injected to prevent hypersensitivity reactions caused by boundary jitter. The time interval dimension uses dynamic temporal elastic registration to find the optimal alignment path between textual temporal expressions and temporal peaks and valleys, outputting a temporal alignment consistency score and a time offset. The offset is directly written into subsequent matrices as evidence of temporal differences. The contextual attribute dimension targets slots such as disaster causes, management behaviors, and reproductive periods, using a relational alignment graph to perform subgraph matching between temporal change patterns and textual relation triples, outputting a semantic consistency score with conflict enumeration. Each of the four main dimensions yields a score and several evidence labels, which are then combined to form a multidimensional similarity result. At the same time, the evidence path and source weight corresponding to each score are retained to ensure traceability.

[0099] The overall consistency score is generated by the weighted fusion module. The fusion weights are derived from three types of observable factors: source weights measure the quality level and historical reliability of the text and image sources; coverage measures the sufficiency of evidence for that dimension on the current object; and freshness measures the timeliness of the data. After normalization, the weights are used to weighted aggregate the scores of the four main dimensions. Conflict flags trigger a suppression gate to reduce the interference of single-dimensional anomalies on the overall judgment. When evidence for a certain dimension is missing, the weight for that dimension is automatically reset to zero according to the missing data mask, and the reason for the missing data is recorded. The generated overall consistency score and the corresponding fusion weight vector are output together, providing a basis for threshold comparison and difference analysis.

[0100] The consistency level is obtained through multi-threshold comparison. The system maintains multiple preset threshold ranges, corresponding to four levels: highly consistent, basically consistent, differing, and significantly abnormal. The overall consistency score is fed into the interval discriminator to give the consistency level, and a safety margin from the boundary is output as a confidence indicator. The thresholds are not used with a fixed configuration for a long time, but are linked to historical comparison records. They are slightly calibrated based on the label statistics and drift detection results of the most recent window. The calibration process is written to the change log and a parameter version number is generated to prevent the thresholds from becoming invalid due to environmental changes.

[0101] The difference analysis matrix is ​​generated after the grading is determined. Structurally, rows represent difference types, and columns record evidence and location information. The type difference row aggregates error pairs and nearest neighbor pairs in the object category dimension, and the columns provide the top-level category pairs, similarity scores, and evidence paths. The area difference row aggregates overlap gaps, spillover ratios, and boundary consistency scores in the spatial coverage dimension, and the columns mark the spatial index and period index of the difference region. The time difference row aggregates time offsets, alignment confidence, and key turning point mismatch indicators, and the columns provide the intervals after parsing the text time expression and their correspondence with time series peaks and troughs. The difference analysis matrix retains the source weight and freshness of each row, facilitating subsequent filtering and interpretation under different business rules.

[0102] The first analysis report is generated based on the combined input of the difference analysis matrix and the consistency level. The report structure includes a comparison conclusion area, a difference summary area, and a handling recommendation area. The comparison conclusion area directly references the consistency level and overall consistency score, and provides the safety margin and parameter version number. The difference summary area presents the location, evidence, and confidence levels of the three types of differences in a tabular format. The handling recommendation area generates actionable recommendations based on the combination of difference types and level mappings. For example, if differences exist and the area differences are concentrated in recent periods, it is recommended to add observations or trigger field verification; if there are significant anomalies and the time differences show a sudden occurrence, it is recommended to initiate the risk management process. The report is output in two tracks: structured data and readable text. The structured portion is used for automatic backfilling of archives and the rule engine, while the readable text is used for manual review and archiving.

[0103] The correlation analysis of historical comparison records is performed immediately after the report is generated. The system retrieves matching historical records using object identifiers and period indexes, calculates the rank distribution and difference pattern stability of the same object within a continuous window, compares the current multidimensional similarity results with historical distributions to identify long-term biases or short-term drifts, and updates the baseline parameters for consistency comparison based on the comparison results, including fine-tuning of the fusion weights of each dimension, minor shifts in threshold boundaries, and alarm sensitivity coefficients for difference types. The update process uses a versioning strategy to write back, retaining old versions to support playback and retrospection, while generating change summaries for auditing purposes.

[0104] This embodiment achieves a unified characterization of four types of evidence—category, space, time, and context—by decomposing and calculating multidimensional similarity results and weighted fusion based on source, coverage, and freshness, and by comprehensively considering the consistency score. This reduces the impact of single-dimensional noise on the overall judgment. Through multi-threshold intervals and drift-aware calibration, the consistency level maintains a stable response to changes in the external environment. Through the structured expression of the difference analysis matrix, type differences, area differences, and time differences are accurately located and can be directly used for automatic decision-making and manual review. By performing correlation analysis between the first analysis report and historical comparison records and updating the baseline parameters for consistency comparison, the system forms an adaptive comparison baseline and parameter convergence effect during continuous operation. This makes the discrimination boundaries and suggestions of subsequent comparisons more consistent with real-world scenarios, thereby improving comparison accuracy and decision reliability while maintaining interpretability.

[0105] In one embodiment, step S50 above includes: S501 divides the target area into several plot units by image segmentation and administrative division matching, and assigns a unique identifier to each plot unit. S502, Based on the time-series feature vector of the target object, a growth curve feature vector is generated for each land parcel unit; S503, establish the mapping relationship between the unique identifier of each plot unit and the corresponding growth curve feature vector; S504, Generate a dynamic status profile of the target object based on the mapping relationship; S505, associate the consistency comparison results in the first analysis report with the growth curve feature vector of the corresponding plot unit, and store the associated results in the target object dynamic status archive. S506, Add timestamp version control information to the dynamic status file of the target object.

[0106] In this embodiment, when dividing the target area into several plot units based on image segmentation and administrative division matching, firstly, semantic segmentation and instance segmentation are jointly inferred on the temporal image to output a land cover category mask and candidate plot boundaries. Then, using administrative divisions and plot boundary vectors as geometric constraints, shape similarity matching and topological correction are performed to remove candidates that cross administrative boundaries and polygon self-intersections, resulting in a topologically closed set of plot units. To ensure consistency in subsequent inter-period tracking, a boundary hash and spatial index key are calculated for each plot unit, and a unique identifier is assigned. The unique identifier consists of a region code, a time slice code, and a boundary hash. Normalization rules prevent the same plot from generating new identifiers under slight boundary jitter. The unique identifier is written into the spatial index and relation index. The spatial index is used for fast location and adjacent retrieval, and the relation index is used for cross-table association within the archive.

[0107] When generating growth curve feature vectors for each land parcel based on the target object's temporal feature vector, period-by-period feature segments are extracted from the target object's temporal feature vector according to the spatial range of the land parcel. Time-axis resampling and missing data compensation are used to form equally spaced sequences. Confidence aggregation is then performed on multi-source observations of the same time slice to obtain a numerical sequence and event-labeled sequence that evolve over time. To enhance inter-period comparability, amplitude normalization, seasonal positional alignment, and outlier suppression are applied to the sequences. Structured summaries such as change point locations, peak and trough locations, and growth slopes are retained and encoded into growth curve feature vectors. The growth curve feature vector comprises four parts: a time index, a multi-dimensional feature trajectory, a structured summary, and a quality label. The time index is used for alignment, and the quality label records the impact of cloud cover, incident angle, and other factors on reliability.

[0108] When establishing the mapping relationship between unique identifiers of each plot unit and the corresponding growth curve feature vectors, a one-to-one primary key mapping is created in the relationship index, and one-to-many version attachments are allowed in the time dimension to record curve history. Before the mapping relationship is written to the archive, it undergoes spatial consistency and temporal continuity checks: spatial consistency checks whether the boundary hash of the unique identifier is consistent with the current boundary, and temporal continuity checks for phase drift and anomalous gaps in adjacent period curves. After the checks pass, the mapping relationship is saved as a primary key record, along with batch and source tags for easy traceability.

[0109] When generating dynamic state archives of target objects based on mapping relationships, an object-oriented archive data model is created. The core layer includes an object metadata layer, a temporal feature layer, a comparison conclusion layer, and a version management layer. The object metadata layer stores unique identifiers, spatial index keys, and static attributes; the temporal feature layer attaches growth curve feature vectors and their quality labels, and maintains an ordered chain along the timeline; the comparison conclusion layer reserves reference entries for analysis results; the version management layer maintains version numbers, parent version numbers, and change summaries, supporting playback by time or event. Archives are implemented using a hybrid row-column storage method: frequently updated curve segments are managed in columnar blocks to improve scanning efficiency, while infrequently updated metadata and conclusions are managed in row-based records to improve point-to-point search efficiency. After archive generation is complete, reverse references from plot units to archive records are registered in the spatial index to ensure that archives can be directly located from spatial retrieval.

[0110] When associating the consistency comparison results in the first analysis report with the growth curve feature vectors of the corresponding land parcel units, the structured fields of the first analysis report are first parsed to extract the consistency level, difference analysis matrix, and evidence path. Then, based on the unique identifier, the target file and the growth curve feature vector of the corresponding time interval are located. The level, difference type, and location information are linked to the comparison conclusion layer as foreign keys. At the same time, a reference pointer is written to the time-series feature layer to form a bidirectional traceable chain from the curve to the conclusion. To avoid data redundancy, the entire report is saved using object storage references, and only the summary and index fields are stored in the file. The associated results are committed using transactional operations to ensure the consistency of references between the curve and the conclusion, and to roll back to the previous stable version in case of failure.

[0111] When adding timestamp version control information to the dynamic state archive of a target object, a multi-granularity timestamp and versioning strategy is introduced. The multi-granularity timestamp includes write time, logical time, and data period time. Write time is used for auditing, logical time for concurrency conflict resolution, and data period time for business alignment. The versioning strategy uses incremental snapshots and change logs in parallel: incremental snapshots save key milestone versions, and change logs record field-level changes and dependency sources, generating version numbers and parent-child chains, supporting retrieval and replay by version number. Each write triggers consistency verification and fingerprint updates. The fingerprint is a combination of a unique identifier, version number, and a core field hash, used for quickly verifying archive integrity. Version metadata and change summaries are entered into the version management layer, and the version view of the spatial and relational indexes is updated synchronously, ensuring that a data view consistent with the specified version is returned when accessing across modules.

[0112] This embodiment, through the joint design of unique identifiers and spatial indexes, maintains stable referencing of land parcel units during cross-period and cross-source fusion, reducing object drift caused by boundary jitter; through the normalization and structured summary expression of growth curve feature vectors, weak-quality observations in the time dimension are robustly aggregated, and subsequent comparisons and early warnings can directly reuse the unified trajectory; through the hierarchical organization of mapping relationships and archive data models, the dynamic status archives of target objects achieve stable latency and traceability in writing, retrieval, and playback scenarios; by establishing bidirectional references between the first analysis report and the growth curve feature vectors and submitting them in a transactional manner, the analysis conclusions and time-series evidence remain consistent, avoiding information fragmentation; by combining timestamp version control information and incremental snapshots, the archives support fine-grained backtracking and concurrent secure updates, reducing maintenance costs and improving data reliability and audit interpretability in a continuous operating environment.

[0113] In one embodiment, step S70 above includes: S701, when an abnormal event is detected, acquire the first time-series image data before the abnormal event occurs and the second time-series image data after the abnormal event occurs; S702, perform change detection processing on the first time-series image data and the second time-series image data, and determine the change detection index; S703, compare the change detection index with a preset threshold to determine the level of influence; S704, Generate preliminary impact analysis results based on the aforementioned impact level; S705, when the confidence level of the impact level is lower than a preset threshold, a review process is triggered; S706. Based on the preliminary impact analysis results or review results, generate a second analysis report containing the impact level and treatment recommendations.

[0114] In this embodiment, abnormal events in the dynamic monitoring link are marked as time slices or spatial units that need to be checked after triggering thresholds or rules. The sources can be three types: level changes of the consistency comparison engine, curve shape mutation detection, and external early warning signal push. In implementation, an event identifier, spatial location (plot unit or grid window), event timestamp and trigger source label are generated for each early warning, and the processing priority is registered in the event queue.

[0115] The first time-series image data represents the continuous observation sequence before the event, and the second time-series image data represents the continuous observation sequence after the event. Both are extracted from the image warehouse according to the spatial range of the object and the time window. The time window is centered on the event timestamp and extends forward and backward by a fixed length or an adaptive length (automatically extended according to cloud cover, revisit cycle and crop phenology). During extraction, item-level quality screening, cloud and shadow masking, geometric fine registration and radiometric normalization are performed. Geometric fine registration uses sub-pixel registration and plot boundary constraints to reduce spurious changes, and radiometric normalization uses multi-temporal relative normalization and observation angle correction to ensure intertemporal comparability. The change detection processing performs difference and comparison operations on two time series after aligning them with the same spatial resolution and temporal rhythm. The processing unit is a set of pixels of the plot unit or a stable raster window, supporting multi-scale pyramids to take into account both details and overall shape. In the optical channel, spectral index difference, spectral angle difference, principal component change vector and texture statistical change are calculated. In the radar channel, backscattering ratio and coherence attenuation are calculated. In the time domain, the growth slope change, peak and valley displacement, inflection point number difference and morphological similarity are calculated. All individual measures are uniformly mapped to dimensionless values ​​and output together with the quality weight.

[0116] The change detection index is a structured encapsulation of the aforementioned multi-source metrics, including index name, statistical region, time pair, statistical method (mean, median, quantile, histogram), confidence label, and outlier contribution. During implementation, it records the change items corresponding to different sensors and bands using expandable fields and saves them in vector format for easy subsequent fusion and backtracking. Preset thresholds are assembled into threshold files based on object type and phenological stage. Threshold sources include historical stable period distribution, expert calibration, and cross-regional migration calibration. During loading, specific files are selected based on region, season, and crop type, and scene correction items can be overlaid (e.g., relaxing optical channel weights during periods of heavy rainfall and emphasizing shortwave infrared items during drought periods). Threshold comparison is performed in two levels: single-item judgment and fusion judgment. Single-item scoring is performed first, and the fusion stage uses a weighted or learning-based fusion strategy to output the degree of influence.

[0117] The impact severity is categorized into multiple levels to characterize the combined intensity of change magnitude and spatial coverage. The calculation considers index magnitude, anomalous pixel proportion, and spatial connectivity simultaneously, providing evidence fields for the level boundaries (the three most contributing indicators and their values). Preliminary impact analysis results, after obtaining the levels, generate structured conclusions including object identifiers, levels, main change types, spatial distribution summaries, and time intervals, along with an evidence snapshot index and quality summary. Main change types are determined through a combination of indicators; for example, a significant decrease in greenness and an increase in the disaster index tend to indicate biomass damage, while increased texture roughness and enhanced backscattering tend to indicate waterlogging or landslides. The output is expressed using labels and parameterized fields instead of natural language descriptions.

[0118] The confidence level quantifies the credibility of preliminary conclusions and is derived from four factors: input image quality, registration residuals, index variance, and model uncertainty. These factors are mapped to values ​​between zero and one using a calibration function. When the confidence level falls below a preset threshold, a review process is triggered. This review process involves an automated sequence of supplementary data and alternative algorithms, without manual intervention. This includes extending the time window to incorporate nearby temporal phases, switching alternative registration algorithms, activating backup sensor channels, changing the fusion strategy, and recalculating indices and levels. The review record retains the differences before and after for auditing purposes. The second analysis report generates a unified structured object after the review is completed. Fields include impact level, main change type, spatial coverage, time interval, key indicator summary, quality and confidence level, whether a review occurred and the difference before and after the review, evidence index, and version number. The report uses immutable storage records and writes them to the conclusion layer of the object archive, while simultaneously establishing pointers on the timeline to support playback and tracing.

[0119] Example Description: In an agricultural insurance fintech scenario, the system first acquires a multi-source dataset of the target area, including satellite remote sensing imagery, drone aerial imagery, land parcel declaration information, disaster reports, meteorological observations, soil moisture, and geographic zoning information. All data is aligned and fused according to time index and spatial coordinates, unifying the time series dimension and spatial raster dimension within a standard fusion framework to form standard fused data. This fusion process not only eliminates resolution and temporal differences between different sensors but also ensures that each data point has a unified reference in the spatiotemporal domain during subsequent calculations, thus providing a quantifiable data foundation for financial decisions such as agricultural insurance claims.

[0120] After obtaining the standard fused data, the system initiates the visual processing module to analyze the image portion. Image features are extracted using multi-scale convolutional layers to extract spatial structure features, which are then processed by the temporal attention mechanism of the Transformer encoder to generate a feature sequence reflecting the changes of land parcels over time. This sequence is then encoded into a temporal feature vector for the target object. This feature vector reflects the spectral changes, coverage area, and texture features of crops at different growth stages, thus supporting subsequent dynamic monitoring and risk identification. Simultaneously, the text processing module performs semantic parsing on the descriptive text portion, extracting structured semantic features from land parcel application materials and disaster reports. Deep semantic modeling is performed on the text using a language model, and knowledge enhancement is achieved by combining it with a domain knowledge base. This ensures that the extracted semantics not only include crop category, land parcel area, and time range, but also identify key information such as disaster cause and damage type. The enhanced semantic features are mapped to a high-dimensional semantic vector space to generate textual features for the target object, forming a comparable feature representation with the image data.

[0121] The system then performs a consistency comparison between the temporal feature vector of the target object and the textual features of the target object. First, it calculates the similarity values ​​of the two in four dimensions: category, space, time, and context, outputting a multi-dimensional similarity result. The results of each dimension are weighted and fused to obtain a comprehensive consistency score, which is then compared with multiple preset thresholds to determine the consistency level, categorized into four levels: highly consistent, basically consistent, with differences, and significantly abnormal. The system simultaneously generates a difference analysis matrix, quantifying type differences, area differences, and time differences, and outputs a first analysis report based on these results. The report not only includes the consistency level and difference data but also includes comparison conclusions and suggestions, providing a reference for subsequent risk control or claims approval. The generated analysis results are correlated with historical comparison records to update the baseline parameters for consistency comparison, enabling the model to continuously self-correct during long-term operation.

[0122] After the consistency comparison is completed, the system constructs a dynamic status profile of the target object using its temporal feature vector. First, through image segmentation and administrative division matching, the target area is divided into multiple plot units, and a unique identifier is assigned to each plot. The temporal features of each plot are extracted as growth curve feature vectors, describing its spectral changes, coverage changes, and health index changes. The system establishes a mapping relationship between the unique identifier and the growth curve feature vector, and generates a dynamic status profile containing spatial, temporal, feature, and historical comparison records. The consistency comparison results from the first analysis report are written into the profile and associated with the growth curve of the corresponding plot, thus ensuring that the profile not only contains status records but also a chain of decision-making evidence. All profile writing operations are accompanied by timestamp version control information to ensure data traceability and multi-version auditing capabilities.

[0123] The system performs real-time monitoring of the dynamic status archives of target objects to identify abnormal events. The monitoring module tracks fluctuations and trends in growth curve characteristics over time and assesses the clustering and persistence of abnormal plots spatially. When the curve shape deviates from historical stable patterns or an abnormal distribution appears in the difference matrix, the system automatically generates an abnormal event record, including event identifier, plot extent, timestamp, and triggering cause. This continuous real-time monitoring process enables the detection of potential risk changes in disaster-stricken areas in the early stages of a disaster, providing advance information for agricultural insurance fund allocation and risk mitigation.

[0124] Upon detecting an anomaly, the system initiates change detection processing. First, it extracts two time-series image data segments before and after the event. After radiometric normalization and geometric registration, it calculates multi-source change detection indicators, including spectral index differences, backscattering attenuation, texture statistical changes, and time-series slope changes. The system compares these indicators with preset thresholds to determine the impact level. The level classification is based not only on indicator values ​​but also on spatial coverage and duration, ensuring the judgment results are both accurate and interpretable. Preliminary impact analysis results are then generated, including change type, coverage area, and time period information. When the confidence level falls below the preset threshold, the system automatically triggers a review process, verifying the results by extending the time window, supplementing satellite sources, or changing the detection algorithm. Finally, a second analysis report is generated, summarizing the impact level, change type, time period, and recommendations, and is linked to the archive to form a complete analysis loop.

[0125] Through the above continuous operations, the system has achieved a fully automated closed loop in the agricultural insurance fintech scenario, encompassing data collection, spatiotemporal fusion, joint understanding of images and text, consistency comparison, dynamic file management, and anomaly identification. The growth status and abnormal changes of each plot are quantified and recorded, with the first and second analysis reports reflecting normal consistency and abnormal changes, respectively. This system, through multi-source data fusion and dynamic update mechanisms, transforms agricultural insurance underwriting and claims processing from traditional manual judgment to traceable, quantifiable, and verifiable intelligent management, providing a data-driven risk control foundation for agricultural insurance financial services.

[0126] This embodiment achieves comparability of change detection indicators in both time and space dimensions through bidirectional time window extraction, quality screening, and geometric and radiometric consistency for abnormal events. By unifying the encapsulation of multi-source, multi-dimensional metrics and conditionally assembling threshold levels, the impact level maintains a stable scale across different seasons and objects. An automated review process driven by confidence introduces supplementary time phases and backup channels for self-correction, reducing the interference of occasional noise on the conclusions. Through structured second analysis reports and archive links, change evidence, levels, and time-series trajectories form a closed-loop expression, allowing subsequent decision-making modules to directly consume unified fields for coordinated processing. This enhances the accuracy, robustness, and traceability of change identification in large-scale, continuous operation environments.

[0127] In one embodiment, a multi-source data spatiotemporal fusion analysis device is provided, which corresponds one-to-one with the multi-source data spatiotemporal fusion analysis method described in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the multi-source data spatiotemporal fusion analysis device of the present invention. The modules include a multi-source data fusion module 10, a visual feature extraction module 20, a text feature extraction module 30, a feature comparison and analysis module 40, a status profile construction module 50, an anomaly monitoring module 60, and a change detection module 70. Detailed descriptions of each functional module are as follows: Multi-source data fusion module 10 is used to acquire multi-source data including time-series image data and descriptive text in the target area, perform spatiotemporal fusion processing on the multi-source data, and obtain standard fused data; The visual feature extraction module 20 is used to analyze the standard fusion data using the visual processing module, extract temporal features related to the image, and obtain the temporal feature vector of the target object. The text feature extraction module 30 is used to analyze the standard fusion data using the text processing module, extract text-related features, and obtain the text features of the target object. The feature comparison and analysis module 40 is used to perform a consistency comparison between the temporal feature vector of the target object and the text features of the target object, and generate a first analysis report. The status profile construction module 50 is used to construct a dynamic status profile of the target object based on the temporal feature vector of the target object, and write the first analysis report into the dynamic status profile of the target object. Anomaly monitoring module 60 is used to monitor the dynamic status file of the target object in real time to identify abnormal events; The change detection module 70 is used to acquire time-series image data before and after the occurrence of the abnormal event and perform change detection processing when an abnormal event is detected, and generate a second analysis report.

[0128] In one embodiment, the multi-source data fusion module 10 is specifically used for: Collect time-series image data of the target area, including satellite remote sensing image data and UAV aerial image data; Collect descriptive text including land parcel declaration information and disaster reports from the target area; Collect IoT sensor data, including meteorological and soil data, from the target area; Collect geographic information data, including administrative divisions and land parcel boundaries, within the target area; Time alignment processing is performed on the time-series image data, the descriptive text, the IoT sensor data, and the geographic information data to obtain time-aligned data; The time-aligned data is spatially aligned to obtain spatially aligned data; Extract the spatiotemporal features of the spatial alignment data to obtain spatiotemporal feature data; Standard fusion data is generated based on the aforementioned spatiotemporal feature data.

[0129] In one embodiment, the visual feature extraction module 20 is specifically used for: Image feature information is decoded from the standard fused data; The image feature information is reconstructed into a feature image sequence; The feature image sequence is processed using multi-scale convolutional layers in the vision processing module to extract multi-scale spatial features and obtain a spatial feature map; The spatial feature map is input into the Transformer encoder of the vision processing module, and the temporal attention mechanism is applied to process the spatial feature map to obtain a temporal feature sequence. Based on the temporal feature sequence, a temporal feature vector of the target object is generated.

[0130] In one embodiment, the text feature extraction module 30 is specifically used for: Semantic information of the text is parsed from the standard fused data; The text semantic information is input into the language model of the text processing module, and the language model is used to perform semantic parsing on the text semantic information to obtain preliminary semantic features. The preliminary semantic features are enhanced by performing knowledge enhancement processing based on a preset domain knowledge base to obtain enhanced semantic features; The enhanced semantic features are mapped to the semantic vector space to generate text features of the target object.

[0131] In one embodiment, the feature comparison and analysis module 40 is specifically used for: Determine the similarity values ​​between the temporal feature vector of the target object and the textual features of the target object in multiple feature dimensions to obtain a multidimensional similarity result; The multidimensional similarity results are weighted and fused to obtain a comprehensive consistency score; The overall consistency score is compared with multiple preset thresholds to determine the consistency level; A difference analysis matrix is ​​generated based on the consistency level, including the analysis results of type differences, area differences, and time differences; Based on the difference analysis matrix and the consistency level, a first analysis report is generated, which includes detailed comparison conclusions, anomaly level analysis, and handling suggestions. The first analysis report is correlated with historical comparison records to update the baseline parameters for consistency comparison.

[0132] In one embodiment, the state profile construction module 50 is specifically used for: By segmenting images and matching administrative divisions, the target area is divided into several plot units, and a unique identifier is assigned to each plot unit. Based on the time-series feature vector of the target object, a growth curve feature vector is generated for each land parcel unit; Establish a mapping relationship between the unique identifier of each plot unit and the corresponding growth curve feature vector; A dynamic status profile of the target object is generated based on the mapping relationship; The consistency comparison results in the first analysis report are associated with the growth curve feature vector of the corresponding plot unit, and the associated results are stored in the dynamic status archive of the target object. Add timestamp version control information to the dynamic status profile of the target object.

[0133] In one embodiment, the change detection module 70 is specifically used for: When an abnormal event is detected, first time-series image data before the abnormal event occurs and second time-series image data after the abnormal event occurs are acquired. Change detection processing is performed on the first time-series image data and the second time-series image data to determine change detection indicators; The change detection index is compared with a preset threshold to determine the level of impact. Preliminary impact analysis results are generated based on the aforementioned impact level. When the confidence level of the impact level is lower than a preset threshold, a review process is triggered; Based on the preliminary impact analysis results or the review results, a second analysis report is generated, which includes the impact level and treatment recommendations.

[0134] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a multi-source data spatiotemporal fusion analysis method on the server side.

[0135] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a multi-source data spatiotemporal fusion analysis method on the client side.

[0136] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps: Acquire multi-source data, including time-series image data and descriptive text, in the target area, and perform spatiotemporal fusion processing on the multi-source data to obtain standard fused data; The standard fusion data is analyzed using the visual processing module to extract temporal features related to the images and obtain the temporal feature vector of the target object. The standard fusion data is analyzed using a text processing module to extract text-related features and obtain the text features of the target object. A consistency comparison is performed between the temporal feature vector of the target object and the textual feature of the target object to generate a first analysis report; A dynamic state profile of the target object is constructed based on the temporal feature vector of the target object, and the first analysis report is written into the dynamic state profile of the target object. Real-time monitoring of the target object's dynamic status profile to identify abnormal events; When an abnormal event is detected, time-series image data before and after the abnormal event is acquired and change detection processing is performed to generate a second analysis report.

[0137] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor: Acquire multi-source data, including time-series image data and descriptive text, in the target area, and perform spatiotemporal fusion processing on the multi-source data to obtain standard fused data; The standard fusion data is analyzed using the visual processing module to extract temporal features related to the images and obtain the temporal feature vector of the target object. The standard fusion data is analyzed using a text processing module to extract text-related features and obtain the text features of the target object. A consistency comparison is performed between the temporal feature vector of the target object and the textual feature of the target object to generate a first analysis report; A dynamic state profile of the target object is constructed based on the temporal feature vector of the target object, and the first analysis report is written into the dynamic state profile of the target object. Real-time monitoring of the target object's dynamic status profile to identify abnormal events; When an abnormal event is detected, time-series image data before and after the abnormal event is acquired and change detection processing is performed to generate a second analysis report.

[0138] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0139] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0140] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0141] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for spatio-temporal fusion analysis of multi-source data, characterized in that, The method comprises the following steps: Obtain multi-source data including time-series image data and description text in a target area, perform spatio-temporal fusion processing on the multi-source data, and obtain standard fusion data; Analyze the standard fusion data using a visual processing module, extract time-series features related to images, and obtain a target object time-series feature vector; Analyze the standard fusion data using a text processing module, extract features related to text, and obtain a target object text feature; Perform consistency comparison on the target object time-series feature vector and the target object text feature, and generate a first analysis report; Construct a target object dynamic state archive based on the target object time-series feature vector, and write the first analysis report into the target object dynamic state archive; Real-time monitor the target object dynamic state archive to identify abnormal events; When an abnormal event is identified, obtain time-series image data before and after the abnormal event and perform change detection processing to generate a second analysis report.

2. The multi-source data spatio-temporal fusion analysis method of claim 1, wherein, Obtain multi-source data including time-series image data and description text in a target area, perform spatio-temporal fusion processing on the multi-source data, and obtain standard fusion data, comprising: Collect time-series image data including satellite remote sensing image data and unmanned aerial vehicle aerial image data in the target area; Collect description text including land declaration information and disaster report in the target area; Collect Internet of Things sensor data including meteorological data and soil data in the target area; Collect geographic information data including administrative division and land boundary in the target area; Perform time alignment processing on the time-series image data, the description text, the Internet of Things sensor data, and the geographic information data to obtain time-aligned data; Perform spatial alignment processing on the time-aligned data to obtain spatially aligned data; Extract spatio-temporal features of the spatially aligned data to obtain spatio-temporal feature data; Generate standard fusion data based on the spatio-temporal feature data.

3. The multi-source data spatio-temporal fusion analysis method of claim 1, wherein, Analyze the standard fusion data using a visual processing module, extract time-series features related to images, and obtain a target object time-series feature vector, comprising: Decode image feature information from the standard fusion data; Reconstruct the image feature information into a feature image sequence; Process the feature image sequence using a multi-scale convolution layer in the visual processing module to extract multi-scale spatial features and obtain a spatial feature map; Input the spatial feature map into a Transformer encoder of the visual processing module, apply a time-series attention mechanism to process the spatial feature map, and obtain a time-series feature sequence; Generate a target object time-series feature vector based on the time-series feature sequence.

4. The multi-source data spatio-temporal fusion analysis method of claim 1, wherein, Analyze the standard fusion data using a text processing module, extract features related to text, and obtain a target object text feature, comprising: Parse text semantic information from the standard fusion data; Input the text semantic information into a language model of the text processing module, use the language model to perform semantic analysis on the text semantic information, and obtain preliminary semantic features; Perform knowledge enhancement processing on the preliminary semantic features based on a pre-set domain knowledge base to obtain enhanced semantic features; Map the enhanced semantic features to a semantic vector space to generate target object text features.

5. The multi-source data spatio-temporal fusion analysis method of claim 1, wherein, Conduct consistency comparison on the target object time sequence feature vector and the target object text features to generate a first analysis report, including: Determine the similarity values of the target object time sequence feature vector and the target object text features in multiple feature dimensions to obtain a multi-dimensional similarity result; Perform weighted fusion processing on the multi-dimensional similarity result to obtain a comprehensive consistency score; Compare the comprehensive consistency score with multiple preset threshold values to determine a consistency level; Generate a difference analysis matrix of the analysis results including type difference, area difference, and time difference according to the consistency level; Based on the difference analysis matrix and the consistency level, generate a first analysis report including detailed comparison conclusions, abnormal level analysis, and processing suggestions; Conduct correlation analysis on the first analysis report and historical comparison records to update the baseline parameters of consistency comparison.

6. The multi-source data spatio-temporal fusion analysis method of claim 1, wherein, Based on the target object time sequence feature vector, construct a target object dynamic state archive, and write the first analysis report into the target object dynamic state archive, including: Divide the target area into several land units through image segmentation and administrative division matching, and assign a unique identifier to each land unit; Generate a growth curve feature vector for each land unit based on the target object time sequence feature vector; Establish a mapping relationship between the unique identifier of each land unit and the corresponding growth curve feature vector; Generate a target object dynamic state archive based on the mapping relationship; Associate the consistency comparison results in the first analysis report with the growth curve feature vector of the corresponding land unit, and store the associated results in the target object dynamic state archive; Add time stamp version control information to the target object dynamic state archive.

7. The multi-source data spatio-temporal fusion analysis method of claim 1, wherein, When an abnormal event is identified, obtain the time sequence image data before and after the abnormal event and perform change detection processing to generate a second analysis report, including: When an abnormal event is identified, obtain the first time sequence image data before the abnormal event and the second time sequence image data after the abnormal event; Perform change detection processing on the first time sequence image data and the second time sequence image data to determine a change detection indicator; Compare the change detection indicator with a preset threshold value to determine an impact level; Generate a preliminary impact analysis result based on the impact level; When the confidence level of the impact level is lower than a preset threshold value, trigger a review process; Generate a second analysis report including the impact level and processing suggestions according to the preliminary impact analysis result or review result.

8. A multi-source data spatio-temporal fusion analysis device, characterized in that, The multi-source data spatio-temporal fusion analysis device includes: A multi-source data fusion module for obtaining multi-source data including time sequence image data and description text in a target area, and performing spatio-temporal fusion processing on the multi-source data to obtain standard fusion data; A visual feature extraction module for analyzing the standard fusion data using a visual processing module to extract time sequence features related to images and obtain a target object time sequence feature vector; The text feature extraction module is configured to analyze the standard fusion data by using the text processing module, extract text-related features, and obtain target object text features. The feature comparison and analysis module is configured to compare and analyze the target object time sequence feature vector and the target object text features for consistency, and generate a first analysis report. The state archive construction module is configured to construct a target object dynamic state archive based on the target object time sequence feature vector, and write the first analysis report into the target object dynamic state archive. The anomaly monitoring module is configured to monitor the target object dynamic state archive in real time to identify abnormal events. The change detection module is configured to, when an abnormal event is identified, acquire time sequence image data before and after the abnormal event and perform change detection processing, and generate a second analysis report.

9. A computer device, comprising: The computer device includes a memory, a processor, and a multi-source data space-time fusion analysis program stored on the memory and executable on the processor. When the multi-source data space-time fusion analysis program is executed by the processor, the steps of the multi-source data space-time fusion analysis method according to any one of claims 1-7 are implemented.

10. A computer-readable storage medium, characterized in that, The storage medium stores a multi-source data space-time fusion analysis program. When the multi-source data space-time fusion analysis program is executed by the processor, the steps of the multi-source data space-time fusion analysis method according to any one of claims 1-7 are implemented.