Geospatial data report generation method and system based on multi-modal information fusion

By using multimodal information fusion and large language models to generate intelligent text content, combined with automated assembly of report templates, the problem of traditional GIS results being limited in expression and lacking interactivity has been solved. This has enabled the generation of efficient and structured multi-format geospatial data reports, thereby enhancing the value of data utilization.

CN121350041BActive Publication Date: 2026-08-25SHANDONG NORMAL UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511936037.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-08-25
Estimated Expiration
2045-12-22

AI Technical Summary

Technical Problem

Traditional GIS deliverables have limited expressive capabilities, are difficult to structure systematically, have poor interactivity, low automation, rely heavily on manual writing for reports, and are difficult to track in terms of dissemination and effectiveness after distribution, making it difficult to realize the value of the data.

Method used

By fusing multimodal information, embedded features are extracted from multi-source heterogeneous geospatial data to generate multimodal input vectors. Intelligent text content is generated using a large language model that is fine-tuned in multiple stages. The content is then automatically assembled using predefined or custom report templates to generate structured reports that support multiple formats.

Benefits of technology

It enhances the expressiveness, dissemination, and feedback capabilities of geospatial data results, achieves systematization, structuring, and interactivity, reduces human intervention, adapts to different scenario needs, supports multi-format output, and enhances the value of data utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121350041B_ABST
    Figure CN121350041B_ABST
Patent Text Reader

Abstract

The present application provides a kind of based on multi-modal information fusion geographical space data report generation method and system, belong to data processing field.Extract embedded feature from original multi-source heterogeneous geographical space related data, and fuse into multi-modal input vector;Embedded feature includes chart structure metadata feature, spatial layer abstract feature, time series statistical abstract feature, table statistical feature and user semantic context prompt feature;The multi-modal input vector is input into the multi-stage fine-tuning multi-modal large language model, and the intelligent text content containing index statement, data interpretation, cause inference and conclusion suggestion is generated;Based on the report template configured by predefinition or user self-definition, the intelligent text content and the corresponding geographical visualization element are automatically assembled to generate a structured report supporting multiple formats.Thereby realize the multi-modal fusion generation of geographical space data report, improve the expression, interaction and automation level of report, improve the application value of geographical space data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method and system for generating geospatial data reports based on multimodal information fusion. Background Technology

[0002] Geospatial data is a digital representation of various geographical features and phenomena on the Earth's surface. With the rapid development of remote sensing technology, the Internet of Things, and big data analytics, its scale and dimensions have grown dramatically. Geographic Information Systems (GIS), as the core tool for processing and analyzing this type of data, leverages its powerful spatial analysis capabilities and is widely used in key areas such as natural resource surveys and ecological environmental protection. It drives the transformation of geospatial knowledge into "archivable, disseminable, and traceable" knowledge products (such as thematic reports), serving scenarios such as results archiving and government transparency.

[0003] However, the efficient delivery of geospatial knowledge faces real obstacles. Traditional GIS deliverables are mostly presented as static maps, scattered statistical charts, or professional internal reports. The complex spatial analysis process and conclusions are difficult for non-experts to understand, limiting the realization of data value in multiple scenarios.

[0004] Existing GIS software also has significant shortcomings in its cartographic output: limited expressive capabilities, monotonous report formats, weak integration of text and graphics, and difficulty in presenting the analysis process and conclusions in a systematic and structured manner; the output PDFs or images lose geographic reference information, and users cannot interact with them, resulting in information loss due to dimensionality reduction; low automation, with report writing and typesetting relying heavily on manual labor, leading to low efficiency and difficulty in standardization; and a lack of user feedback, making it difficult to track the dissemination scope and effectiveness of the results after distribution, creating a "black box of data usage" that is not conducive to assessing the true value of data products. Summary of the Invention

[0005] To address the aforementioned issues, this invention proposes a method and system for generating geospatial data reports based on multimodal information fusion, which comprehensively enhances the expressiveness, dissemination, and feedback capabilities of geospatial data results.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a method for generating geospatial data reports based on multimodal information fusion, comprising: Embedded features are extracted from the original multi-source heterogeneous geospatial data and fused into a multimodal input vector. The embedded features include chart structure metadata features, spatial layer summary features, time series statistical summary features, table statistical features, and user semantic context prompt features. The multimodal input vectors are fed into a multimodal large language model that has undergone multi-stage fine-tuning to generate intelligent text content that includes indicator statements, data interpretation, causal inferences, and conclusion suggestions. Based on predefined or user-defined report templates, the intelligent text content is automatically assembled with corresponding geographic visualization elements to generate structured reports that support multiple formats.

[0007] Secondly, the present invention provides a geospatial data report generation system based on multimodal information fusion, comprising: The multimodal input acquisition module is configured to extract embedded features from the original multi-source heterogeneous geospatial data and fuse them into a multimodal input vector; the embedded features include chart structure metadata features, spatial layer summary features, time series statistical summary features, tabular statistical features, and user semantic context prompt features; The text generation module is configured to input multimodal input vectors into a multimodal large language model that has undergone multi-stage fine-tuning, and generate intelligent text content that includes indicator statements, data interpretations, causal inferences, and conclusion suggestions. The report generation module is configured to automatically assemble the intelligent text content with the corresponding geographic visualization elements based on a predefined or user-defined report template, and generate a structured report that supports multiple formats.

[0008] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the geospatial data report generation method based on multimodal information fusion described in the first aspect.

[0009] Fourthly, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps in the geospatial data report generation method based on multimodal information fusion described in the first aspect.

[0010] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention utilizes multimodal information fusion to extract rich embedded features from multi-source heterogeneous geographic data and generate input vectors. A large language model, fine-tuned through multiple stages, generates intelligent text containing multi-dimensional content. Combined with flexible predefined or custom report templates, the text and geographic visualization elements are automatically assembled. This solves the problems of traditional GIS outputs being monotonous in expression and lacking interactivity, making reports systematic, structured, and interactive. It also improves automation, reduces manual intervention, and allows templates to adapt to different scenario needs. Furthermore, it lays the foundation for subsequent feedback and tracking of output usage, enhancing the utilization value of geospatial data.

[0011] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0012] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute a limitation thereof.

[0013] Figure 1 The main flowchart of a geospatial data report generation method based on multimodal information fusion provided in this embodiment of the invention is shown below. Detailed Implementation

[0014] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0015] Example 1 like Figure 1 As shown in the figure, this embodiment discloses a method for generating geospatial data reports based on multimodal information fusion, including the following steps: S1: Extract embedding features from the original multi-source heterogeneous geospatial data and fuse them into a multimodal input vector; the embedding features include chart structure metadata features, spatial layer summary features, time series statistical summary features, table statistical features, and user semantic context prompt features; S2: Input the multimodal input vector into the multimodal large language model that has been fine-tuned in multiple stages to generate intelligent text content that includes indicator statements, data interpretation, causal inference and conclusion suggestions; S3: Based on predefined or user-defined report templates, the intelligent text content is automatically assembled with the corresponding geographic visualization elements to generate structured reports that support multiple formats.

[0016] Next, combined Figure 1 This embodiment provides a detailed description of a geospatial data report generation method based on multimodal information fusion.

[0017] I. Data Modeling and Fusion Acquire multi-source heterogeneous geospatial data and business attribute data, and perform unified organization, standardized expression, and semantic alignment.

[0018] Geospatial data includes spatial layers and remote sensing images. Spatial layers are vector and raster data that represent geographic locations and their attributes. They are mainly used to display geographic information such as urban planning, land use, and transportation networks. They usually exist in the form of points, lines, and areas and carry related spatial attribute data. Remote sensing images are high-resolution image data acquired through remote sensing technologies such as satellites or drones. They are mainly used to monitor large-scale changes in the Earth's surface, such as vegetation cover, climate change, and disaster monitoring.

[0019] Business attribute data includes attribute indicator tables and service call logs. Attribute indicator tables contain various statistical data and business information related to geospatial data, and are typically used to record and display characteristics such as the quantity, area, type, and classification of various geographic regions or spatial objects. Service call logs record log data generated when users or systems call geographic information services, including information such as call time, request content, response result, and execution status, and are used to monitor, analyze, and optimize service performance.

[0020] Multiple data sources, including spatial layers, remote sensing imagery, attribute index tables, and service call logs, are formally represented as a set. : (1) in, Represents a collection of spatial vector layers (such as administrative boundaries and land use distribution). Represents a sequence of remote sensing images. This represents an attribute statistics table or an observation index matrix. This represents service call and access log data.

[0021] First, all this data is merged into a unified geographic dataset built on PostGIS. To ensure consistency in the fusion, a geometric registration operator is performed in the spatial dimension: (2) A collection of spatial vector layers representing a unified spatial reference frame after geometric registration; This represents a spatial geometric registration operator that maps data with different resolutions and projection systems to a unified spatial reference frame based on geographic coordinate system unification and spatial overlay operations.

[0022] Subsequently, through spatial registration function and time resampling function By projecting all the data onto a unified spatiotemporal grid system, a standardized spatiotemporal modal tensor is obtained. : (3) In the formula, Represents spatial coordinates; Use time coordinates; It serves as a modality identifier to distinguish different modality types, such as spatial layer modality, remote sensing image modality, attribute index table modality, or service log modality; A collection of spatial regions For time sets, This represents the number of modal channels.

[0023] For each spatial unit Multimodal features It can be represented as: (4) in, They represent the first Spatial vector layer features, remote sensing image features, attribute index features, and service call and access log features of each spatial unit. To achieve semantic alignment and information enhancement, the system employs a modal embedding function. Projecting different modalities into a unified semantic space: (5) in, For learnable parameters, This is a non-linear activation function. The process can be viewed as a composite mapping of multimodal semantic compression and feature alignment. The contributions of different modalities are determined through an adaptive weighted fusion strategy. (6) in, Indicates the first Multimodal fusion semantic features of spatial units This represents the total number of modes, and T represents the transpose operation. For modal attention weights, For modal query vectors, This represents the semantic feature vector obtained by the embedding mapping function for the i-th spatial unit in the m-th modality. This represents the query feature or reference modality embedding of the i-th spatial unit in the k-th modality. This mechanism enables the system to adaptively select the most contributing modality according to different task scenarios, achieving dynamic multimodal balance at the semantic level. Finally, the fusion tensor can be formally defined as: (7) in For embedded dimensions, This represents the total number of spatial units. This represents the total number of time units. The standardization process is achieved by minimizing the modal alignment error: (8) in, It serves as a modal semantic similarity mask, used to constrain the consistent representation of geographical entities with the same origin in multimodal space; This represents the embedding vector of the i-th spatial unit in the m-th mode. Let represent the embedding vector of the j-th spatial unit in the n-th modality, used to represent the individual feature representation of the entity in the semantic space of that modality. Through backpropagation and iterative optimization, this process ensures semantic consistency and scale uniformity among different modalities.

[0024] In summary, through four steps—spatial registration, temporal resampling, feature embedding, and attention fusion—a high-dimensional, multimodal, and differentiable geographic data standardization framework was constructed. This framework not only achieves unified organization of heterogeneous data but also lays a standardized foundation for subsequent intelligent content generation and semantic reasoning. This process ensures that vector boundaries, raster pixel values, and business logs can all be understood and processed by the model in a unified manner.

[0025] II. Intelligent Content Generation The core of content generation based on a multimodal large language model is a specially fine-tuned multi-task multimodal large language model. This model is built on the advanced LLAMA 3 13B parameter structure, and its unique feature lies in its powerful multimodal information understanding capability.

[0026] The input for fine-tuning LLAMA is not a single text cue, but rather a fusion of five different types of multimodal embedding vectors. Specifically, from the fusion tensor... Five types of features are extracted: chart structure metadata, spatial layer summary features, time series statistical summary, table statistical results, and user-provided semantic context prompts, to construct a multimodal input vector.

[0027] The chart structure metadata is extracted by parsing the chart type, axis information, data series, legend labels, and color mapping to form a structural description of the chart. These features, along with spatial layers, time-series statistics, and tabular data, provide LLAMA with a rich, cross-modal information representation, enabling it to effectively process various types of data and make corresponding inferences and predictions. The multimodal input vector representation integrating five types of features is as follows: : (9) in, This indicates a feature fusion operation. Represents multimodal input vector The dimensions. These five types of embedding features. These represent the metadata of the chart structure. Spatial layer summary features Time Series Statistical Summary Statistical results in tables and user-provided semantic context hints For example, time series statistical summaries It includes vectors describing temporal evolution features such as periodic mean, standard deviation, and gradient of change; while spatial layer summary features This includes spatial pattern information such as layer area, number of patches, and proportion of dominant land types: (10) (11) In equation (10), This represents a feature vector for time series statistical summaries, containing the periodic mean of the time series. Standard deviation First-order rate of change Second-order rate of change and peak time These statistical characteristics describe the trend of time evolution and are used to reflect the dynamic changes of indicators over time.

[0028] In equation (11), Represents the feature vector of a spatial layer summary, containing the area of ​​the layer or region. Number of plaques The proportion of dominant land types Layer level or ribbon label and land category codes Elements describing spatial patterns are used to characterize the structural distribution features of geographical entities.

[0029] Multimodal input vector After processing through a multi-head attention mechanism and a feedforward network layer, the data is uniformly input into the Transformer backbone network. The model encodes this information to obtain the hidden state tensor. And finally generate intelligent text content through a Softmax layer. : (12) In the formula, This is the weight matrix. This is the bias term. The model is based on a preset prompt template: indicator statement, data description, causal inference, and conclusion suggestion. Its output strictly follows the complete processing flow from data source indicator to data description, and then to the generation of backpropagation.

[0030] It is worth noting that the model's output strictly follows the process from indicator statement to data explanation, then to causal inference, and finally to conclusions and suggestions, ensuring the logic and professionalism of the generated text.

[0031] To ensure that each generated section of content is traceable, source path meta-annotations are attached to each paragraph during the content generation process. These source path meta-annotations not only provide contextual support for the generated text but also indicate the source data and processing procedures, ensuring clear traceability for every part of the report.

[0032] Suppose a certain generated paragraph Based on a certain dataset a certain feature The generated source path meta-annotation can be represented as: (13) in, It is the data source identifier (such as the ID of a remote sensing image or statistical table). These are features extracted from the data source (such as climate change or farmland change trends in a certain region). It refers to data processing procedures (such as data normalization, spatiotemporal registration, etc.).

[0033] By adding source path meta-annotations to each generated paragraph All generated content is traceable and verifiable. In the formula... middle, This represents the model's predicted output. The model represents the inference path fused from multiple processing modules. Source path meta-annotations, by labeling and tracking each step of the inference path, ensure that the source of data flow and information flow are clearly recorded. This makes the input and output of each inference step not just the computational result, but a finely tuned process, ensuring that the role and data source of each link are clearly labeled, enhancing the model's interpretability and professionalism.

[0034] Through meticulous annotation and labeling, we ensure that the generated results conform to the expected logical structure and meet specific professional standards. Each step of the reasoning process involves not only processing the input data but also reasoning and integrating the relationships between data sources, ensuring that the final output is reasonable, reliable, and possesses a high level of accuracy.

[0035] The generated text is not only logically consistent, but also clearly shows how each data point affects the final result, avoiding potential errors or biases, and making the entire model operation process more transparent, which facilitates subsequent analysis and verification.

[0036] III. Structural Processing (a) Generation of structured reports Structured report generation is a core component of this process. It is responsible for automatically formatting geographic analysis results into multi-page, multi-figure, and multi-language reports according to structured document standards.

[0037] Based on predefined report templates, intelligent text content is automatically assembled with corresponding geographic visualization elements to generate structured reports that support multiple formats.

[0038] This embodiment uses JasperReport as the underlying engine, combined with template-based development, to support complex rules for mixed text and graphics. It should be understood that JasperReport leverages this technology to provide flexible report design, supports dynamic presentation of different data formats, and achieves precise positioning and high-quality rendering of geographic visualization elements through pixel-level precision control.

[0039] During the generation of structured reports, intelligent text wrapping and context-aware automatic pagination enable intelligent processing of the relationship between text and graphics. This allows for automatic adjustment of pagination logic and overall layout based on dynamic changes in geographic data, and also ensures that the arrangement of text and graphics is more closely aligned with the content, improving the visual harmony of the presentation.

[0040] Meanwhile, a template-based data binding mechanism is adopted, and geospatial metadata is embedded into the report structure using semantic markup. This design ensures both the professional visual presentation of the report and its machine-readable structure, providing fundamental support for subsequent automated processing and in-depth data analysis.

[0041] In terms of multilingual adaptation, it supports dynamic language adjustment of content and can automatically switch language versions according to the needs of different regions. It is especially optimized for bilingual scenarios in Chinese and English, and can intelligently adapt to the page width and character length of different languages. It adopts right-hand layout rules to ensure the standardization of text display, which greatly improves the applicability and flexibility of the report globally.

[0042] In terms of output formats, the system is compatible with multiple types such as PDF, Word, HTML, GeoPDF, and Markdown, which can meet different publishing needs such as archiving and storage, printing and distribution, and web page embedding. Among them, the GeoPDF format supports deep integration of map layers. By preserving geographic reference coordinates, readers can perform interactive operations such as map zooming, feature clicking, and spatial positioning within the document, opening up a new path for "interactive archiving" of spatial results.

[0043] It also comes equipped with a template library covering different application scenarios, allowing users to choose according to their actual needs, such as: 1. Special Topic Analysis Report: Applicable to spatial analysis on a specific topic (such as arable land quality or meteorological anomalies), highlighting the analysis area, methods, conclusions, and recommendations.

[0044] 2. Trend Assessment Report: Shows the changing trends of spatial indicators over time, covering monthly, quarterly, and annual dimensions.

[0045] 3. Administrative Division Comparison Report: Used for horizontal comparative analysis between multiple administrative units, functional zones, etc., combining maps, tables and comparative conclusions for output.

[0046] 4. Annual Change Report: Reflects the evolution of the same object across multiple data periods, and supports functions such as overlay comparison and linked chart expression.

[0047] (ii) Custom configuration and interactive report generation In addition to predefined report templates, users can customize configurations. Customization and interactive reporting are geared towards end users, providing a high degree of freedom in report generation, with a particular emphasis on human-computer interaction and interpretability of results. Users can customize the following elements through the system's visual interface: 1. Analysis object selection: Supports configuration by various geographic objects such as task area, administrative area, ecological function area, land use unit, etc. The system automatically aggregates the corresponding data and graphics.

[0048] 2. Statistical dimension settings: Users can select time granularity (year, quarter, month), thematic dimension (resource type, environmental element), spatial level (province, city, county, township), etc.

[0049] 3. Chart style configuration: The system has built-in various chart types such as line chart, bar chart, stacked chart, radar chart, pie chart, and heat map. Users can drag and drop to combine them and preview the effect in real time.

[0050] 4. Chapter customization function: Supports users to input personalized titles, abstracts, conclusions, and suggested paragraphs. The system provides enhanced functions such as large language model-assisted polishing and data-driven conclusion generation.

[0051] Furthermore, map layers support interactive displays. For example, users can click on a region in a chart to highlight the region's map outline and overlay a heatmap with statistical data; or click on a region on the map to dynamically update the table and chart content on the right side of the report, achieving a "what you see is what you get" data exploration experience. This submodule highly integrates spatial visualization technology Mapbox, a web component framework, and a reporting engine, deeply integrating GIS representation with document compilation.

[0052] In this embodiment, predefined templates unify report formats and data standards, ensuring professionalism and consistency across different scenarios, eliminating repetitive design steps, and improving report generation efficiency. User-defined configurations, on the other hand, can adapt to personalized business needs, flexibly adjusting data dimensions, display formats, etc., to meet the analysis and decision-making needs of different scenarios. The combination of the two avoids the tediousness and errors of manual production, and allows structured reports to quickly respond to routine needs while adapting to special scenarios.

[0053] (III) Report Assembly and Output The final stage of report generation is automated assembly and document output. This stage employs an assembly mechanism based on template-driven and modal references to achieve unified organization and formatted output of various analysis results. The system manages multimodal content such as chapters, charts, tables, and text through a Document Object Model (DOM), and drives the generation of the final document based on a predefined set of chapters. (14) In this formula, The final generated report document consists of multiple components. Generated through rendering. This represents the rendering function. Each component... Includes the following three elements: This is the title of a section in the report, usually a brief description of the content of that section. This refers to the way or format in which this part is presented, such as text, charts, or data tables. This refers to the specific content of this section, namely the actual data or information displayed in relation to the title and modality.

[0054] These elements (title, modality, and content) are combined and processed according to specific rendering rules. The title identifies the theme or category of each section, the modality determines the presentation of the content (such as text, charts, or maps), and the content is the specific data or information. Through this processing method, the title, modality, and content of each section are effectively organized and rendered, ensuring a clear report structure, accurate information expression, and the ability to integrate multiple presentation formats, such as text, charts, maps, and other multimedia elements, ultimately generating a complete and highly interactive report.

[0055] The report generation process goes beyond simply merging titles, modalities, and content to create a static document. To further enhance the report's interactivity and functionality, the system integrates geographic information into the report, generating GeoPDF format files with native interactive features. In this process, the system utilizes the Adobe Acrobat SDK and the MapPublisher plugin, allowing data from each spatial layer to be embedded into the report, ensuring that native interactivity is preserved in the PDF report, such as zooming in and out of the map and viewing detailed information.

[0056] (15) In this formula, This represents the geographic information layer at each level in the report. This represents the two-dimensional spatial coordinates of the layer. This includes metadata related to the layer (such as coordinate system, projection information, etc.). This spatial data and metadata are embedded in the final generated PDF report, allowing users to interact directly with the map when viewing the report.

[0057] In addition to traditional text and data presentation, the generated report will automatically produce statistical reports and analytical modules for analysis, and further enhance the presentation of spatial data through the GeoPDF format. Through these steps, the generated report is not only interactive but also provides users with a more dynamic and richer information experience.

[0058] IV. Multi-stage fine-tuning 1. LLAMA fine-tuning based on geographic entity data To enable LLAMA models to have basic understanding and processing capabilities in the geographic domain, fine-tuning based on geographic entity data is used to allow general LLAMA models to fully understand geospatial and geographic features.

[0059] This fine-tuning process employs an architecture that combines a pre-trained language model with retrieval enhancements, leveraging external geographic knowledge bases to improve model performance.

[0060] (1) Data preparation and search library construction First, various types of geographic entity data are collected, such as administrative divisions, natural geographic features, man-made facilities, and annual change trends of geographic entities, and processed into text format. Simultaneously, a retrieval knowledge base is constructed based on historical and predictive data of geographic entities. Vectorization techniques are used to convert this textual information into high-dimensional vectors, and an index is created for rapid retrieval.

[0061] (2) Query and Retrieval Enter text related to geographic information (such as questions or generated tasks) and generate query text.

[0062] The query text is encoded into a vector, and the most relevant geographic entity documents are retrieved by calculating the similarity (such as cosine similarity) with document vectors in the database.

[0063] (3) Model fine-tuning Enter text That is, text related to geographic information and retrieved documents containing relevant geographic entities. They are concatenated together as input to the model.

[0064] For example, the input text consists of geographic entities (such as "Jinan") and retrieved related geographic entity documents (such as "Jinan's climate type, Jinan's recent climate changes, Jinan's transportation, mountains, water bodies, etc."). The concatenation of the input can be represented as follows: (16) here, This indicates the separator between the input and retrieved documents. Indicates input text The word segmentation or sub-token sequence, This indicates the retrieved relevant geographic entity documents. The word segmentation or sub-token sequence.

[0065] The process involves forward propagation, where the input query text and related documents are encoded into corresponding representation vectors by the encoder. These vectors are then passed to the decoder for further processing via a self-attention mechanism and a multi-layer neural network.

[0066] The model aims to generate a geographic description relevant to the input query. Assuming the target output is y, for example: "Jinan's climate is a temperate monsoon climate." The model generates a predicted output based on the input query and retrieved geographic entity documents. The loss function includes: a. Language model loss (autoregressive loss) Similar to standard language models, language model loss calculates the difference between the model-generated text and the actual target output. The loss calculation formula is as follows: (17) in, It is the length of the target output. It is the first of the target outputs One word, The model represents the i-th word. The predicted probability.

[0067] b. Search enhancement loss To enhance the model's geographical knowledge, an augmentation loss is retrieved. This is used to optimize the relevance of retrieved documents. The retrieval loss can be calculated by determining the similarity between the input query and the retrieved document. (18) in, It is the embedding vector of the query text. It is the embedding vector of the retrieved document. This represents a similarity metric.

[0068] c. Total Loss Function The total loss function of the fine-tuning process is a weighted sum of the language model loss and the retrieval augmentation loss: (19) in, It is a hyperparameter used to control the trade-off between language model loss and retrieval loss.

[0069] Furthermore, backpropagation and parameter updates are performed. The gradient is calculated using the backpropagation algorithm, and the model parameters are updated using an optimization algorithm (such as Adam). The update rule is as follows: (20) in, These are the current model parameters. It's the learning rate. It is the gradient of the total loss function with respect to the model parameters.

[0070] (4) Verification and evaluation During fine-tuning, the model performance is evaluated periodically using a validation set containing known answers.

[0071] The evaluation metrics mainly include the accuracy of the generated content and the relevance of the search results.

[0072] Through this series of processes, the fine-tuned LLAMA model can gain a deeper understanding and generate geo-related professional content. When dealing with complex geographical tasks, it can effectively utilize external knowledge to provide more accurate reasoning and answers.

[0073] 2. LLAMA Fine-tuning Based on Behavioral Feature Vectors To further enable the model to generate more accurate distribution reports and overcome the shortcoming of traditional GIS results being "invisible to users" after distribution, this embodiment designs a distribution business statistics support submodule.

[0074] This module connects to external business system logs to collect real-time usage behavior of distributed reports. This module connects to external business systems (such as project management platforms, data authorization platforms, and government data transfer platforms), and the data collected includes the following dimensions: (1) Project name, task ID, and user organization; (2) The field to which the project belongs (natural resources, agriculture and rural areas, ecological and environmental protection, etc.); (3) Data usage behavior: access frequency, download records, number of API calls; (4) Data service types: thematic map services, data download, data analysis API, etc.; (5) Purpose of use: research and analysis, rule evaluation, results presentation, decision support, etc.; (6) Service coverage: The spatial range and accuracy level of the data.

[0075] The system automatically generates reports from this statistical information, including bar charts (number of projects or frequency of use), heatmaps (regional distribution), and trend line charts (time-series evolution of usage behavior), supplemented with text annotations and table outputs. Through this module, managers can quantitatively evaluate the service value, distribution efficiency, and coverage of different data products, improving the closed-loop capability of application results.

[0076] These behavioral data cover multiple dimensions of business interactions, and a high-dimensional behavioral tensor is constructed based on these dimensions. It fully depicts the entire picture of the use of the results after they are distributed: (twenty one) in It is a matrix containing data related to behavioral characteristics. Representing the set of real numbers, the dimensions of the matrix are determined by... , ,..., The matrix is ​​composed of dimensions representing different behaviors or characteristics. Each dimension corresponds to a specific variable, such as task type, user behavior type, time interval, or spatial distribution. Therefore, each element in the matrix represents the relationship or influence between different tasks and behavioral characteristics. In this way, the formula describes how behavioral characteristics are combined and calculated across multiple dimensions to help analyze the impact of tasks on behavior, thereby optimizing the system's performance.

[0077] Considering the high-dimensional sparsity of this tensor, the system first projects it into the embedding space, and then uses methods such as truncated singular value decomposition to reduce its dimensionality, resulting in a compact behavioral feature vector. : (twenty two) in, This indicates the truncated singular value decomposition operation. This indicates an embedded function.

[0078] This behavioral feature vector is correlated with domain information in external business system logs. User intent After being fused with other context vectors, they together constitute the input to the LLAMA model for generating the distribution report and fine-tuning it. : (twenty three) Before model decoding, a dynamic high-density threshold determination mechanism was introduced in the distribution business statistics support submodule to achieve intelligent recognition of access behavior. This mechanism utilizes a behavior tensor constructed from access logs. Statistical analysis was performed to calculate the access density of each spatial unit within a preset spatiotemporal window. And based on historical average visits with standard deviation Dynamically determine the threshold: (twenty four) Among them, parameters The system adaptively adjusts based on access stability. When the current access density meets... > At that time, the system determines that the spatial unit is in a high-density access state. To enhance robustness, the system can also combine a machine learning model based on temporal anomaly detection to dynamically adjust the statistical threshold, ultimately generating a fusion threshold: (25) Wherein, β is the fusion coefficient, which controls the weight balance between the two; This represents a static statistical threshold obtained based on historical access log statistics. It is used to reflect the average level and fluctuation range of access density under long-term stable conditions. Its value can be obtained by linear combination of the access density mean and standard deviation. The dynamic calculation results, as part of the statistics submodule, are integrated into... This enables a unified framework for multi-level threshold fusion. This represents the judgment threshold dynamically calculated on the real-time access behavior vector by anomaly detection or machine learning models, which can adaptively reflect short-term mutations and abnormal patterns in access behavior.

[0079] Based on the above-mentioned judgment mechanism, the model can identify high-density clustering patterns of access behaviors in the spatiotemporal domain and generate analysis paragraphs about usage behaviors by combining domain context information. : (26) in, This indicates that the LLAMA output Z is decoded. For example, it generates data-driven conclusions such as "In the third quarter of 2024, the frequency of calls to farmland remote sensing products increased by 42.6% quarter-on-quarter..."

[0080] The semantic structure of a paragraph is mapped by a quadruple. Control was implemented to ensure the entity Numerical values ,time and spatial location Accurate expression: (27) In this embodiment, during the fine-tuning stage based on geographic entity data, building a retrieval knowledge base and generating query text by inputting commands allows the model to initially grasp basic concepts and entity relationships in the geographic domain, laying the foundation for subsequent optimization. Fine-tuning based on behavioral feature vectors, by invoking behavioral feature vectors, domain information, and user intent, allows the model to further optimize the generated content according to the user's actual usage habits and preferences, making it more aligned with real-world application scenarios and improving the model's accuracy and practicality in geospatial data report generation tasks.

[0081] This specific embodiment utilizes multimodal information fusion to extract rich embedded features from raw, multi-source, heterogeneous geospatial data and fuse them into multimodal input vectors, significantly enhancing the comprehensiveness and depth of data utilization. A multimodal large language model, fine-tuned through multiple stages, generates intelligent text content covering various aspects such as indicator descriptions, enriching the textual meaning of reports. Automated assembly based on predefined or user-defined report templates not only improves report creation efficiency and reduces manual intervention but also meets diverse user needs. Furthermore, the generated structured reports support multiple formats, enhancing their versatility and dissemination, comprehensively releasing the potential value of geospatial data, and powerfully promoting the efficient application of geospatial knowledge in scenarios such as results archiving, government transparency, and decision analysis.

[0082] Example 2 This embodiment provides a geospatial data report generation system based on multimodal information fusion, including: The multimodal input acquisition module is configured to extract embedded features from the original multi-source heterogeneous geospatial data and fuse them into a multimodal input vector; the embedded features include chart structure metadata features, spatial layer summary features, time series statistical summary features, tabular statistical features, and user semantic context prompt features; The text generation module is configured to input multimodal input vectors into a multimodal large language model that has undergone multi-stage fine-tuning, and generate intelligent text content that includes indicator statements, data interpretations, causal inferences, and conclusion suggestions. The report generation module is configured to automatically assemble the intelligent text content with the corresponding geographic visualization elements based on a predefined or user-defined report template, and generate a structured report that supports multiple formats.

[0083] Example 3 This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in the geospatial data report generation method based on multimodal information fusion as described in Embodiment 1 above.

[0084] Example 4 This embodiment provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps in the geospatial data report generation method based on multimodal information fusion as described in Embodiment 1 above.

[0085] The steps or modules involved in Embodiments 2 to 4 above correspond to those in Embodiment 1. For specific implementation details, please refer to the relevant description section of Embodiment 1. The term "computer-readable storage medium" should be understood as a single medium or multiple media including one or more instruction sets; it should also be understood as including any medium capable of storing, encoding, or carrying an instruction set for execution by a processor and enabling the processor to perform any of the methods in this invention.

[0086] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for generating geospatial data reports based on multimodal information fusion, characterized in that, include: Embedded features are extracted from the original multi-source heterogeneous geospatial data and fused into a multimodal input vector; The embedded features include chart structure metadata features, spatial layer summary features, time series statistical summary features, table statistical features, and user semantic context prompt features; Multimodal input vectors are input into a multimodal large language model that has undergone multi-stage fine-tuning to generate intelligent text content containing indicator statements, data interpretations, causal inferences, and conclusion suggestions. The multi-stage fine-tuning includes fine-tuning based on geographic entity data and fine-tuning based on behavioral feature vectors. Fine-tuning based on behavioral feature vectors includes: using user reports to fine-tune the model using features; the user reports using features include behavioral feature vectors, domain information, and user intent; the behavioral feature vectors... Specifically: ; ; in, This indicates the truncated singular value decomposition operation. Indicates an embedded function; It is a matrix containing data related to behavioral characteristics. Represents the set of real numbers; the dimensions of the matrix are determined by... , ,..., The matrix is ​​composed of dimensions representing different behaviors or characteristics; each dimension corresponds to a variable, including task type, user behavior type, time interval or spatial distribution. Each element in the matrix represents the relationship or influence between different tasks and behavioral characteristics, which helps to analyze the impact of tasks on behavior. In the process of generating intelligent text content, source path meta-annotations are introduced to mark and record the details of the reasoning path at each step; specifically, it is assumed that a generated paragraph is based on a certain dataset. a certain feature The generated source path meta-annotation is represented as: in, It is a data source identifier; These are features extracted from the data source; It is a data processing process; Add source path meta-annotations to each generated paragraph. ,in, This represents the model's predicted output. This represents the inference path after being integrated through multiple processing modules; the source path meta-annotation ensures that the source of the data flow and the flow of information are clearly recorded by marking and tracking each step of the inference path; each step of inference is not only the processing of input data, but also includes the inference and integration of relationships between data sources. Based on predefined or user-defined report templates, the intelligent text content is automatically assembled with corresponding geographic visualization elements to generate structured reports that support multiple formats; specifically, geographic information is integrated into the report: ; in, This represents the geographic information layer at each level in the report. This represents the two-dimensional spatial coordinates of the layer. This is metadata associated with the layer.

2. The method for generating geospatial data reports based on multimodal information fusion as described in claim 1, characterized in that, The original multi-source heterogeneous geospatial data includes geospatial data and business attribute data; The geospatial data includes spatial layers and remote sensing images; The business attribute data includes attribute indicator tables and service call logs.

3. The method for generating geospatial data reports based on multimodal information fusion as described in claim 1, characterized in that, The time-series statistical summary features are vectors describing the characteristics of time evolution, including periodic mean, standard deviation, and gradient of change; the spatial layer summary features include spatial pattern information such as layer area, number of patches, and proportion of dominant land types.

4. The method for generating geospatial data reports based on multimodal information fusion as described in claim 1, characterized in that, The fine-tuning based on geographic entity data includes: A retrieval knowledge base is built based on pre-collected geographic entity data, and query text is generated by inputting commands; The instruction text and query text are concatenated to obtain the input vector; The input vector is fed into a multimodal large language model, and the model is fine-tuned through forward propagation and backward propagation based on the preset target output.

5. A geospatial data report generation system based on multimodal information fusion, comprising the geospatial data report generation method based on multimodal information fusion as described in claim 1, characterized in that, include: The multimodal input acquisition module is configured to extract embedded features from the original multi-source heterogeneous geospatial related data and fuse them into a multimodal input vector. The embedded features include chart structure metadata features, spatial layer summary features, time series statistical summary features, table statistical features, and user semantic context prompt features; The text generation module is configured to input multimodal input vectors into a multimodal large language model that has undergone multi-stage fine-tuning, and generate intelligent text content that includes indicator statements, data interpretations, causal inferences, and conclusion suggestions. The report generation module is configured to automatically assemble the intelligent text content with the corresponding geographic visualization elements based on a predefined or user-defined report template, and generate a structured report that supports multiple formats.

6. The geospatial data report generation system based on multimodal information fusion as described in claim 5, characterized in that, The original multi-source heterogeneous geospatial data includes geospatial data and business attribute data; The geospatial data includes spatial layers and remote sensing images; The business attribute data includes attribute indicator tables and service call logs.

7. The geospatial data report generation system based on multimodal information fusion as described in claim 5, characterized in that, The time-series statistical summary features are vectors describing the characteristics of time evolution, including periodic mean, standard deviation, and gradient of change; the spatial layer summary features include spatial pattern information such as layer area, number of patches, and proportion of dominant land types.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps in the geospatial data report generation method based on multimodal information fusion as described in any one of claims 1-4.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps in the geospatial data report generation method based on multimodal information fusion as described in any one of claims 1-4.

Citation Information

Patent Citations

  • Intelligent report generation method and system based on multi-modal fusion

    CN121052236A