Geographic information element analysis method and system based on multi-modal large model

By using a multimodal large-scale model for geographic information element analysis, dynamically adjusting modal weights and combining a three-level verification mechanism, the problem of insufficient data fusion accuracy in existing technologies is solved, achieving more accurate and comprehensive analysis results.

CN120910171AInactive Publication Date: 2025-11-07HUNAN CHUANGXIN WEILI TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510977335.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-16
Publication Date
2025-11-07
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In geographic information analysis, existing technologies cannot dynamically adjust the priority of key data when fusing multimodal data, resulting in insufficient analysis accuracy and easy information gaps or biases.

Method used

A geographic information element analysis method based on a multimodal large model is adopted. By integrating satellite remote sensing imagery, ground GNSS positioning data, meteorological station observation texts and UAV oblique photography videos, a dual-driven attention mechanism of "goal orientation - modal contribution" is designed to dynamically adjust modal weights and fuse them. Combined with a three-level verification mechanism, the analysis accuracy is ensured.

Benefits of technology

It significantly improves the flexibility and targeting of the analysis process, enhances the accuracy of capturing target geographical phenomena, reduces information gaps, and provides more comprehensive and reliable analysis results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120910171A_ABST
    Figure CN120910171A_ABST
Patent Text Reader

Abstract

The invention discloses a geographic information element analysis method and system based on a multi-modal large model. According to the method, the flexibility and pertinence of the analysis process are remarkably improved through a dynamic self-adaptive multi-modal fusion mechanism. Through a task sensing module and real-time contribution degree calculation, key data types in different scenes can be automatically identified, the dynamic adjustment enables an analysis process to better meet actual requirements, interference of irrelevant data is avoided, the capture precision of a target geographic phenomenon is directly enhanced, and the accuracy of target geographic phenomenon capture is improved. And the comprehensiveness and the reliability of an analysis result are enhanced through deep collaborative fusion of multi-modal data. Various types of data are complemented and cooperated through unified feature space mapping and dual-drive weight adjustment, not only are multi-dimensional features of geographic information covered, but also limitation of a single data source is reduced through real-time verification and iterative optimization, and a finally output analysis result is closer to a real geographic environment. And a more solid information support is provided for subsequent decision making.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of geographic information systems, and specifically relates to a geographic information element analysis method and system based on a multi-modal large model. BACKGROUND

[0002] Geographic information element analysis is one of the core contents of geographic information systems (GIS), which refers to the process of systematically studying natural and cultural elements on the earth's surface through spatial data collection, processing, modeling, and analysis. The analysis objects include topography, hydrological network, land use, vegetation cover, climate characteristics, population distribution, transportation routes, administrative divisions, and other geographic elements, aiming to reveal their spatial distribution patterns, mutual relationships, and dynamic change characteristics. This analysis process usually combines spatial query, buffer analysis, network analysis, terrain analysis, overlay analysis, and other technical means, and is widely used in urban planning, disaster assessment, environmental protection, resource management, transportation optimization, and public policy making. Geographic information element analysis not only relies on remote sensing, global positioning system (GPS), and geographic database technologies, but also integrates statistical, computer science, and geographic methods, and is an important technical foundation for spatial decision support and smart city construction.

[0003] However, in the prior art, fixed weight fusion of different modal data is often used in geographic information analysis, which cannot dynamically adjust the priority of key data according to specific analysis targets, resulting in insufficient accuracy in capturing target phenomena. At the same time, due to the large differences between different data forms (such as images, texts, and videos), information gaps or one-sidedness may occur in the fusion process, and the limitations of single data source are obvious, resulting in weak comprehensiveness and reliability of the analysis results. SUMMARY

[0004] The purpose of the present application is to solve the above-mentioned problems, and to provide a geographic information element analysis method and system based on a multi-modal large model.

[0005] The technical solution adopted by the present application is as follows: a geographic information element analysis method based on a multi-modal large model, the method comprising the following steps:

[0006] S1: Integrate satellite remote sensing images, ground GNSS positioning data, meteorological station observation texts, unmanned aerial vehicle oblique photography videos, and other multi-modal data, wherein the remote sensing images focus on macroscopic land cover, the GNSS data provide accurate coordinate points, the meteorological texts supplement climate elements, and the unmanned aerial vehicle videos are used for capturing local object details; this step provides the basis for the subsequent preprocessing of the original data, and the collection frequency of each data source needs to be dynamically adjusted according to the target area;

[0007] S2: Radiometric correction, geometric correction of the collected remote sensing image, unified to WGS-84 coordinate system; interpolation of GNSS discrete point data into grid elevation model; extraction of temperature, precipitation and other structured numerical values from meteorological text using natural language processing technology; frame extraction of unmanned aerial vehicle video and labeling of key features through target detection algorithm; the preprocessed data needs to retain the original modal feature identification for subsequent steps to trace the modal source;

[0008] S3: Based on the encoder of the multi-modal large model, the pixel features of the remote sensing image, the spatial coordinate features of the GNSS, the numerical features of the meteorological text, and the time sequence features of the unmanned aerial vehicle video are mapped to the same high-dimensional semantic space;

[0009] S4: Break through the traditional fixed weight fusion mode, design a "target-oriented - modal contribution" double-driven attention mechanism: first, activate the task perception module of the large model according to the analysis target, generate the initial modal weight; At the same time, calculate the contribution of each modal data to the current analysis target in real time, dynamically adjust the weight and fuse;

[0010] S5: Use the fused high-dimensional features to output the geographic element classification results through the decoder of the large model: identify the land cover type, extract the terrain parameters, and label the special features; The feature space information in step S3 needs to be associated during the extraction process;

[0011] S6: Based on the labeled data set of historical geographical events, train the association rule module of the large model, and establish the multi-element influence relationship of "terrain-climate-human activity";

[0012] S7: Adopt a three-level verification mechanism: micro-scale comparison of unmanned aerial vehicle video manual labeling of ground features and model extraction results; mesoscale verification of element value accuracy using ground test points; macro-scale test of regional overall classification accuracy through authoritative geographic information products; If the verification fails, go back to step S4 to adjust the fusion strategy or step S6 to optimize the association rules until the accuracy threshold is met;

[0013] S8: Integrate the verified element extraction results and the association rules into visual products, including hierarchical thematic maps, time sequence change animations, and warning prompts; The real-time data acquisition module in step S1 needs to be associated during the output process. When new data is input, automatically trigger the rapid iteration of steps S4 to S7 to ensure the timeliness of the analysis results.

[0014] In a preferred embodiment, in step S1, the data collection covers four types of core data sources: satellite remote sensing, ground GNSS, weather station observations, and unmanned aerial vehicle oblique photography. Satellite remote sensing uses multispectral and synthetic aperture radar sensors, where multispectral images include visible light, near-infrared, and thermal infrared bands, with a spatial resolution of 1-30 meters and a revisit period adjusted according to the target area. GNSS positioning data is collected through a CORS station network with a sampling frequency of 1 Hz and a positioning accuracy better than 0.1 meters, and the coverage density is set according to the analysis requirements. Weather station observation data includes hourly temperature, precipitation, wind speed, humidity, and other parameters provided by the national weather station network in CSV text format. Unmanned aerial vehicle oblique photography uses multi-rotor unmanned aerial vehicles with a flight height of 100-300 meters, an aerial photography overlap of 60%-80%, a video shooting frame rate of 25 frames per second, and a coverage range adjusted according to the target area.

[0015] In a preferred embodiment, in step S2, remote sensing image preprocessing includes radiation correction and geometric correction: radiation correction uses the absolute calibration method to convert pixel gray values to surface reflectivity based on sensor calibration parameters; geometric correction uses a rational polynomial coefficient model combined with GNSS control points for orthorectification, with an error of less than 1 pixel after correction between the corrected image and the WGS-84 coordinate system; GNSS discrete point data is converted to a gridded elevation model using the inverse distance weighted interpolation method, with a grid resolution matching the remote sensing image.

[0016] In a preferred embodiment, in step S3, a multi-modal encoder with a Transformer architecture is used to map different modal data to a 512-dimensional high-dimensional semantic space; the encoder includes image branches, text branches, and numerical branches, and the outputs of each branch are fused through an attention mechanism; during feature mapping, contrastive learning is used to constrain cross-modal semantic consistency: for different modal data of the same geographic area, the cosine similarity of their feature vectors is calculated, and the InfoNCE loss function is used to maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs.

[0017] In a preferred embodiment, in step S4, first, the analysis target is converted into a task embedding vector t by the task perception module of the large model, which contains the core semantics of geographic information elements: surface temperature, building density, vegetation cover, and other key elements; then, multi-modal data (infrared remote sensing image features f ir , visible light image features f vis , surface temperature text and numerical features f temp , and weather station wind speed features f wind ) are input into the encoder to obtain high-dimensional representations of each modality; the task perception module generates initial weights αi (0)For example, infrared images have a significantly higher initial weight than visible light images because they directly reflect the thermal radiation of the ground surface.

[0018] At the same time, the system calculates the contribution of each modality to the current analysis target in real time: for infrared images, the similarity of their features to the historical thermal island area infrared feature library is calculated; for land surface temperature text, the deviation of its numerical value from the thermal island threshold is calculated; for wind speed data, the inhibitory ability of heat diffusion is calculated; finally, the dynamic weight αi is determined by the initial weight αi (0) combined with the contribution βi, after weight adjustment, the features of each modality are weighted and fused into comprehensive features f fusion for subsequent thermal island area identification and intensity calculation.

[0019] The task-aware initial weight generation formula is:

[0020]

[0021] Where: t is the task embedding vector (dimension d, generated by the large model according to the geographical information element target, containing temperature, underlying surface type); f i is the high-dimensional feature vector of the i-th modality; the denominator is the sum of the exponents of all modalities, ensuring weight normalization;

[0022] The modality contribution calculation formula is:

[0023] β i = γ i · Sim(f i , f i,ref );

[0024] Where: γi is the modality-specific coefficient: infrared image γir=1.5, land surface temperature text γtemp=1.2, visible light image γ vis =0.8, reflecting the prior importance of different modalities to thermal island analysis); Sim(·) is the cosine similarity function, f i,ref is the historical thermal island feature reference vector of the i-th modality (such as the reference vector of the infrared image, which is the average infrared feature of the historical thermal island area);

[0025] Dynamic weight adjustment formula:

[0026]

[0027] Where: is the task-aware initial weight (range 0-1), β i is the modality contribution (range 0-2, constrained by γi and similarity), and the adjusted weight α i dynamically amplifies the influence of high-contribution modalities.

[0028] The multi-modal fusion formula is:

[0029]

[0030] wherein f fusion is the fused comprehensive feature vector (dimension d), a i is the dynamically adjusted modal weight (satisfying ∑ i a i = 1), and f i is each modal high-dimensional feature vector.

[0031] In a preferred embodiment, in the step S5, based on the fused high-dimensional features, an improved version of U-Net decoder is used for element extraction; the decoder input is 512-dimensional fused features, and the output includes three branches: surface cover classification, terrain parameter extraction, and special feature labeling; the classification layer uses a Softmax activation function, the terrain parameter layer uses a regression loss, and the labeling layer uses a Focal Loss.

[0032] In a preferred embodiment, in the step S6, the association rule learning is based on a historical 10-year geographic event labeling dataset, each event is labeled with multi-element information such as terrain parameters, weather data, and surface cover types; the rule learning module uses an LSTM time series model, the input is a multi-element time series, and the output is an event occurrence probability; the model training uses a cross-entropy loss, the optimizer is Adam, the training rounds are 100, and the validation set accounts for 20%.

[0033] In a preferred embodiment, in the step S7, verification is performed at three levels of micro, meso, and macro scales: the micro scale is based on manual labeling of unmanned aerial vehicle video, the IoU of element extraction is calculated, and the IoU of buildings, vegetation, and other features is required to be > 0.7; the meso scale uses ground measurement sample points to verify the accuracy of terrain parameters and numerical elements; the macro scale compares the national basic geographic information database to calculate the overall accuracy of surface cover classification, and the requirement is > 90%.

[0034] In a preferred embodiment, the analysis result in step S8 is output through a visualization platform, including three types of products: hierarchical thematic map, time series change animation, and early warning prompt; the hierarchical thematic map is classified using the natural breakpoint method, and the color mapping uses a red-yellow-blue gradient color system; the time series change animation is generated based on the annual extraction results over a period of 10 years, with a frame rate of 24 frames per second, and each frame corresponds to 1 year of data, visually displaying the processes of urban expansion, vegetation change, etc.; the early warning prompt is triggered by real-time data: when new meteorological data is collected in S1, the system automatically triggers the rapid iteration of S4 to S7, and if the analysis result shows that the heat island intensity exceeds the threshold, the user is pushed to the early warning information; the service output frequency is adjusted according to the data update period: the regular analysis is updated daily, and the early warning analysis is pushed in real time, ensuring the timeliness of the information and the value of the decision support.

[0035] In summary, due to the adoption of the above technical solutions, the present application has the following advantages:

[0036] 1. In the present application, the dynamic and adaptive multi-modal fusion mechanism significantly improves the flexibility and pertinence of the analysis process. Traditional methods often use fixed weight to fuse different data, which is difficult to adjust the focus according to the specific analysis target. However, this method can automatically identify the key data types in different scenarios through the task perception module and real-time contribution calculation, such as focusing more on infrared images and temperature text when analyzing urban heat islands, and increasing the weight of unmanned aerial vehicle video and wind speed data when analyzing forest fire danger. This dynamic adjustment makes the analysis process more in line with actual needs, avoiding the interference of irrelevant data and directly enhancing the capture accuracy of the target geographical phenomenon.

[0037] 2. In the present application, the deep collaborative fusion of multi-modal data enhances the comprehensiveness and reliability of the analysis results. In the past, due to the large differences between different data forms (such as pixels of images, numerical values of texts, and time series of videos), information gaps or one-sidedness often occurred after fusion. However, this method uses unified feature space mapping and dual-driven weight adjustment to make various data truly "complementary and collaborative": remote sensing images provide macroscopic surface features, GNSS data supplement accurate coordinates, meteorological texts improve climate influences, and unmanned aerial vehicle videos refine local details. The effective collaboration of multi-modal data not only covers the multi-dimensional features of geographic information, but also reduces the limitations of single data sources through real-time verification and iterative optimization, ultimately outputting analysis results that are closer to the real geographical environment and providing more solid information support for subsequent decision-making. BRIEF DESCRIPTION OF DRAWINGS

[0038] Figure 1 The figure is a schematic diagram of the flow principle of the present application. DETAILED DESCRIPTION

[0039] In order to make the purpose, technical scheme and advantages of the present application more clear, the present application is further described in detail below in combination with the drawings and examples. It should be understood that the specific examples described herein are only used to explain the present application and not to limit the present application.

[0040] Embodiments:

[0041] With reference to Figure 1 A geographic information element analysis method based on a multi-modal large model, the method comprising the following steps:

[0042] S1: Integrate satellite remote sensing images (including visible light, infrared, radar bands), ground GNSS positioning data, meteorological station observation text, unmanned aerial vehicle oblique photography video and other multi-modal data. The remote sensing images focus on macroscopic land cover, the GNSS data provide accurate coordinate points, the meteorological text supplements climate elements, and the unmanned aerial vehicle video is used for capturing local ground object details. This step provides a raw data basis for subsequent preprocessing and needs to dynamically adjust the data source collection frequency according to the target area (such as mountains or cities).

[0043] S2: Perform radiation correction and geometric correction on the collected remote sensing images and unify them to the WGS-84 coordinate system; interpolate the GNSS discrete point data into a gridded elevation model; extract temperature, precipitation and other structured numerical values from the meteorological text using natural language processing technology; frame the unmanned aerial vehicle video and label key ground objects (such as buildings and vegetation) through a target detection algorithm. The preprocessed data needs to retain the original modal feature identifiers (such as the band information of the image and the timestamp of the text) to trace the modal source in the subsequent steps.

[0044] S3: Based on the encoder of the multi-modal large model, map the pixel features of the remote sensing images, the spatial coordinate features of the GNSS, the numerical features of the meteorological text and the time sequence features of the unmanned aerial vehicle video to the same high-dimensional semantic space. For example, through contrastive learning, map "green patches in the image of a certain area" to "vegetation coverage value of the corresponding GNSS point" and "precipitation in the meteorological text of that month" to similar vectors, solve the semantic fragmentation problem caused by the difference in data form of different modalities and provide a computable unified representation for multi-modal fusion.

[0045] S4: Break through the traditional fixed weight fusion mode and design a "target-oriented - modal contribution" double-driven attention mechanism: first, activate the task perception module of the large model according to the analysis target (such as "urban heat island effect" or "forest fire danger level") to generate initial modal weights; at the same time, calculate the contribution of each modal data to the current analysis target (such as infrared images and ground temperature text have higher contribution than visible light images when analyzing heat island), dynamically adjust the weights and fuse.

[0046] S5: Utilize the fused high-dimensional features to output geographic element classification results through the decoder of the large model: identify land cover types (such as farmland, buildings, water areas), extract terrain parameters (slope, slope direction), and label special features (such as landslide bodies, transportation hubs). In the extraction process, the feature space information in step S3 needs to be associated, for example, the "fusion features of a certain pixel point" are traced back to the original remote sensing bands, GNSS coordinates, to ensure that the extraction results are interpretable.

[0047] S6: Based on the labeled data set of historical geographic events (such as floods, urban expansion), train the association rule module of the large model to establish the multi-element influence relationship of "terrain-climate-human activities". For example, learn the rule "slope > 25° + monthly precipitation > 200mm → landslide probability increases by 40%", and continuously verify and correct the rule parameters through the real-time extraction results of step S5, forming a "data → rule → data" closed-loop optimization.

[0048] S7: Adopt a three-level verification mechanism: micro-scale comparison of unmanned aerial vehicle video manual labeling of ground features and model extraction results; meso-scale verification of element value accuracy using ground measurement samples (such as soil moisture sensor data); macro-scale verification of regional overall classification accuracy through authoritative geographic information products (such as national basic geographic information database). If the verification fails, go back to step S4 to adjust the fusion strategy or step S6 to optimize the association rules until the accuracy threshold (such as overall classification accuracy > 90%) is met.

[0049] S8: Integrate the verified element extraction results and association rules into visual products, including hierarchical thematic maps (such as ecological sensitivity distribution maps), time series change animations (such as urban expansion process over the past ten years), and early warning prompts (such as geological disaster risk level based on current weather data). In the output process, the real-time data acquisition module in step S1 needs to be associated, and when new data (such as heavy rain warning) is input, automatically trigger the rapid iteration of steps S4 to S7 to ensure the timeliness of the analysis results.

[0050] In step S1, data collection covers four types of core data sources: satellite remote sensing, ground GNSS, weather station observations, and unmanned aerial vehicle oblique photography. Satellite remote sensing uses multi-spectral and synthetic aperture radar (SAR) sensors, where multi-spectral images include visible light (400-700 nm), near-infrared (700-1100 nm), and thermal infrared (8-14 μm) bands, with a spatial resolution of 1-30 meters and a revisit period adjusted according to the target area (5 days / visit for urban areas and 2-3 days / visit for mountainous areas). GNSS positioning data is collected through a CORS station network with a sampling frequency of 1 Hz and a positioning accuracy better than 0.1 meters, with a coverage density set according to analysis requirements (3-5 points per square kilometer in urban areas and 1-2 points per square kilometer in mountainous areas). Weather station observation data includes hourly temperature, precipitation, wind speed, humidity, and other parameters provided by the national weather station network in CSV text format. Unmanned aerial vehicle oblique photography uses multi-rotor drones with a flight height of 100-300 meters, an aerial photography overlap of 60%-80%, a video shooting frame rate of 25 frames / second, and a coverage range adjusted according to the target area size. During the collection process, the priority of each data source needs to be dynamically adjusted according to the analysis target: for urban heat island analysis, focus on collecting thermal infrared images and temperature text with high frequency; for geological disaster warning, increase the density of SAR images and GNSS points.

[0051] In step S2, remote sensing image preprocessing includes radiation correction and geometric correction: radiation correction uses the absolute calibration method to convert pixel gray values to surface reflectance based on sensor calibration parameters; geometric correction uses the Rational Polynomial Coefficient (RPC) model combined with GNSS control points (5-10 per image) for orthorectification, with an error of less than 1 pixel in the WGS-84 coordinate system. GNSS discrete point data is converted into a gridded elevation model using the inverse distance weighted interpolation method (IDW), with a grid resolution matching the remote sensing image (e.g., 30-meter resolution images correspond to 30-meter x 30-meter grids). Weather text data uses natural language processing techniques to extract temperature, precipitation, and other numerical values through a named entity recognition model, converting them into a structured table with a timestamp in the "YYYY-MM-DD HH:MM" format. Unmanned aerial vehicle videos are extracted at a rate of 2 frames per second, and a YOLOv8 object detection model is used to label ground features, with the labeling results stored as XML files containing feature categories, coordinate boundaries, and video frame timestamps. The preprocessed data needs to be attached with a modal identification field (e.g., "remote sensing_thermal infrared", "GNSS_elevation", "weather_temperature") for modal tracing in subsequent steps.

[0052] In step S3, a multi-modal encoder with a Transformer architecture is used to map different modal data to a 512-dimensional high-dimensional semantic space. The encoder includes an image branch (ResNet50 feature extraction + Transformer encoder), a text branch (BERT text encoding + linear projection), and a numerical branch (fully connected layer + normalization), and the outputs of each branch are fused through an attention mechanism. During feature mapping, contrastive learning is used to constrain cross-modal semantic consistency: for different modal data of the same geographic area, the cosine similarity of their feature vectors is calculated, and the InfoNCE loss function is used to maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs. After mapping, the feature vectors of each modality in the semantic space are adjacent to each other, ensuring the semantic interpretability of subsequent fusion.

[0053] In step S4, the analysis target is first converted into a task embedding vector t by a task-aware module of a large model, which contains the core semantics of geographic information elements such as surface temperature, building density, and vegetation coverage. Then, multi-modal data (infrared remote sensing image features f ir , visible light image features f vis , surface temperature text and numerical features f temp , and weather station wind speed features f wind ) are input into the encoder to obtain high-dimensional representations of each modality. The task-aware module generates initial weights αi (0) by calculating the semantic correlation between the task embedding and the representations of each modality, for example, the initial weight of the infrared image is significantly higher than that of the visible light image because the infrared image directly reflects the surface thermal radiation.

[0054] At the same time, the system calculates the contribution of each modality to the current analysis target in real time βi: for the infrared image, the similarity between its features and the historical heat island region infrared feature library is calculated; for the surface temperature text, the deviation of its numerical value from the heat island threshold is calculated; for the wind speed data, the inhibition ability of heat diffusion is calculated (the lower the wind speed, the higher the contribution). The final dynamic weight αi is determined by the initial weight αi (0) and the contribution βi. After weight adjustment, the modality features are weighted and fused into a comprehensive feature f fusion , which is used for subsequent heat island region identification and intensity calculation.

[0055] The task-aware initial weight generation formula is as follows:

[0056]

[0057] Where: t is the task embedding vector (dimension d, generated by the large model based on geographic information elements, including temperature and underlying surface type); fi is the high-dimensional feature vector of the i-th modality (dimension d, output by the multimodal encoder); the denominator is the sum of the exponents of all modalities to ensure weight normalization.

[0058] The formula for calculating modal contribution is:

[0059] β i =γ i Sim(f) i ,f i,ref );

[0060] Where: γi is a mode-specific coefficient (e.g., γir = 1.5 for infrared imagery, γtemp = 1.2 for surface temperature text, and γi = 1.2 for visible light imagery). vis =0.8, reflecting the prior importance of different modes to heat island analysis; Sim(·) is the cosine similarity function, f i,ref This is the reference vector for the historical heat island features of the i-th mode (e.g., the reference vector for infrared images is the average infrared feature of the historical heat island region).

[0061] Dynamic weight adjustment formula:

[0062]

[0063] in: For initial weights of task perception (range 0-1), β i The adjusted weight α represents the modal contribution (range 0–2, constrained by γi and similarity). i The influence of high-contribution modes is dynamically amplified.

[0064] The multimodal fusion formula is:

[0065]

[0066] Where: f fusion Let α be the fused composite feature vector (dimension d). i The dynamically adjusted modal weights (satisfying ∑ i α i =1), f i These are the high-dimensional feature vectors for each modality.

[0067] In step S5, based on the fused high-dimensional features, an improved version of the U-Net decoder is used for element extraction. The decoder input is a 512-dimensional fused feature, and the output includes three branches: surface cover classification (10 categories, including farmland, building, water area, forest, etc.), terrain parameter extraction, and special feature labeling. The classification layer uses a Softmax activation function, the terrain parameter layer uses a regression loss, and the labeling layer uses a Focal Loss. During extraction, the feature space information in S3 is associated through an attention mechanism: each extraction result needs to backtrack to the original modality feature, outputting its corresponding thermal infrared image brightness value, GNSS elevation value, and temperature text value, ensuring that the classification result can be traced back to the feature support of at least two modalities, improving reliability.

[0068] In step S6, the association rule learning is based on a historical 10-year geographic event labeling dataset (including 2000 landslide events, 500 urban expansion cases, and 300 heat island extreme records), each event labeled with terrain parameters, weather data, and surface cover types, etc. The rule learning module uses an LSTM time series model, with multi-element time series (such as daily rainfall, daily average temperature, and slope values in the previous 30 days) as input, and event occurrence probability (such as landslide probability and heat island intensity) as output. The model training uses cross-entropy loss, the optimizer is Adam (learning rate 0.001), the training rounds are 100 times, and the validation set accounts for 20%. The rule parameters are optimized through gradient descent. In the iterative optimization phase, when the error between the newly extracted element data and the model predicted event probability exceeds 5%, the parameter correction is triggered: the training set is expanded using new data, the first 5 layers of the model are retrained, the rule parameters are updated, and the adaptability of the rule to real-time data is ensured.

[0069] In step S7, verification is performed at three levels: micro, meso, and macro. The micro scale is based on manual labeling of unmanned aerial vehicle video (labeling accuracy is pixel level, covering 5% of the target area), calculating the IoU (intersection over union) of element extraction, requiring IoU > 0.7 for buildings, vegetation, and other features. The meso scale uses ground measurement points (50 points per 100 square kilometers, including soil moisture, ground temperature, etc.), verifying the accuracy of terrain parameters and numerical elements. The macro scale compares with the national basic geographic information database, calculating the overall accuracy of surface cover classification, requiring > 90%. If any scale fails the verification, it is backtracked and corrected according to priority: if the micro scale fails, adjust the decoder parameters in S5; if the meso scale fails, optimize the fusion weights in S4; if the macro scale fails, retrain the association rule model in S6. After correction, the verification needs to be performed again until all three scales meet the accuracy requirements.

[0070] In step S8, the analysis result is output through a visualization platform, including three types of products: hierarchical thematic map, time series change animation, and early warning prompt. The hierarchical thematic map (such as the heat island intensity map) is classified using the natural breakpoint method (5 levels, corresponding to "none-weak-medium-strong-extremely strong"), and the color mapping uses a red-yellow-blue gradient color system; the time series change animation is generated based on the annual extraction results of the past 10 years, with a frame rate of 24 frames per second, and each frame corresponds to 1 year of data, which intuitively shows the process of urban expansion, vegetation change, etc. The early warning prompt is triggered by real-time data: when new meteorological data is collected in S1, the system automatically triggers the rapid iteration of S4 to S7, and if the analysis result shows that the heat island intensity exceeds the threshold, the early warning information is pushed to the user. The service output frequency is adjusted according to the data update period: the regular analysis is updated daily, and the early warning analysis is pushed in real time, ensuring the timeliness of the information and the value of decision support.

[0071] A geographic information element analysis system based on a multi-modal large model, the system runs the above-mentioned geographic information element analysis method based on a multi-modal large model when in use.

[0072] From the above, it can be seen that:

[0073] In the present application, the dynamic adaptive multi-modal fusion mechanism significantly improves the flexibility and pertinence of the analysis process. Traditional methods often use fixed weight to fuse different data, which is difficult to adjust the focus according to the specific analysis target, while this method can automatically identify the key data types in different scenarios through the task perception module and real-time contribution calculation - for example, more emphasis on infrared images and temperature text when analyzing urban heat islands, and increasing the weight of unmanned aerial vehicle video and wind speed data when analyzing forest fire danger. This dynamic adjustment makes the analysis process more in line with actual needs, avoiding interference from irrelevant data and directly enhancing the capture accuracy of the target geographic phenomenon.

[0074] In the present application, the deep collaborative fusion of multi-modal data enhances the comprehensiveness and reliability of the analysis results. In the past, due to the large differences between different data forms (such as pixels of images, numerical values of texts, and time series of videos), information gaps or one-sidedness often occurred after fusion, while this method uses unified feature space mapping and dual-driven weight adjustment to make various data truly "complementary and collaborative": remote sensing images provide macroscopic surface features, GNSS data supplement accurate coordinates, meteorological texts improve climate influences, and unmanned aerial vehicle videos refine local details. The effective collaboration of multi-modal data not only covers the multi-dimensional features of geographic information, but also reduces the limitations of single data sources through real-time verification and iterative optimization, so that the final output of the analysis result is closer to the real geographic environment, providing more solid information support for subsequent decision-making.

[0075] It is to be noted that, in the present text, relational terms such as first and second and the like can be used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any actual such relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", "has", "having", "includes", "including", or any other variation thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or even inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a", "has... a", "includes... a", or "has... a" does not, without more constraints, preclude the existence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0076] The above examples are merely intended to illustrate the technical solutions of the present application, but not to limit the same; even though the present application has been described in detail with reference to the foregoing examples, it should be understood by those skilled in the art that the technical solutions recorded in the foregoing examples can be modified, or some technical features thereof can be replaced by equivalents; and such modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for geographic information element analysis based on a multi-modal large model, characterized in that: The method comprises the following steps: S1: integrate satellite remote sensing images, ground GNSS positioning data, weather station observation text, unmanned aerial vehicle oblique photography video and other multi-modal data, wherein the remote sensing image focuses on macro land cover, the GNSS data provides accurate coordinate points, the weather text supplements climate elements, and the unmanned aerial vehicle video is used for capturing local ground object details; this step provides the original data basis for subsequent preprocessing, and the collection frequency of each data source needs to be dynamically adjusted according to the target area; S2: the collected remote sensing images are subjected to radiation correction and geometric correction, and are unified to the WGS-84 coordinate system; the GNSS discrete point data is interpolated into a gridded elevation model; the weather text is extracted by using natural language processing technology to extract temperature, precipitation and other structured numerical values; the unmanned aerial vehicle video is frame extracted and the key ground objects are labeled by using a target detection algorithm; the preprocessed data needs to retain the original modal feature identification, so as to trace the modal source in the subsequent steps; S3: based on the encoder of the multi-modal large model, the pixel features of the remote sensing image, the spatial coordinate features of the GNSS, the numerical features of the weather text and the time sequence features of the unmanned aerial vehicle video are mapped to the same high-dimensional semantic space; S4: break through the traditional fixed weight fusion mode, and design a "target-oriented-modal contribution" double-driven attention mechanism: first, activate the task perception module of the large model according to the analysis target, generate the initial modal weight; at the same time, calculate the contribution degree of each modal data to the current analysis target in real time, dynamically adjust the weight and fuse; S5: using the fused high-dimensional features, the decoder of the large model outputs the geographic element classification results: identifying land cover types, extracting terrain parameters, labeling special ground objects; the feature space information of step S3 needs to be associated in the extraction process; S6: based on the labeled data set of historical geographical events, the correlation rule module of the large model is trained, and the multi-element influence relationship of "terrain-climate-human activity" is established; S7: a three-level verification mechanism is adopted: the micro-scale compares the ground object labels manually labeled by the unmanned aerial vehicle video with the model extraction results; the meso-scale verifies the element value accuracy by using the ground measured sample points; the macro-scale verifies the overall classification accuracy of the region by using the authoritative geographic information product; if the verification fails, the fusion strategy of step S4 or the correlation rule of step S6 is adjusted until the accuracy threshold is met; S8: the verified element extraction results and the correlation rules are integrated into visual products, including hierarchical thematic maps, time sequence change animations and early warning prompts; in the output process, the real-time data collection module of step S1 needs to be associated, and when new data is input, the rapid iteration of steps S4 to S7 is automatically triggered, so as to ensure the timeliness of the analysis results.

2. The geographic information element analysis method based on a multi-modal large model according to claim 1, characterized in that: In the step S1, the data collection covers four types of core data sources, including satellite remote sensing, ground GNSS, weather station observation, and unmanned aerial vehicle oblique photography. The satellite remote sensing adopts multispectral and synthetic aperture radar sensors, wherein the multispectral image contains visible light, near-infrared, and thermal infrared bands, with a spatial resolution of 1-30 meters and a revisit period adjusted according to the target area. The GNSS positioning data is collected through the CORS station network, with a sampling frequency of 1 Hz and a positioning accuracy better than 0.1 meters, and the coverage density is set according to the analysis requirements. The weather station observation data includes hourly temperature, precipitation, wind speed, humidity, and other parameters provided by the national weather station network, with a CSV text format. The unmanned aerial vehicle oblique photography uses a multi-rotor unmanned aerial vehicle, with a flight height of 100-300 meters, an aerial photography overlap of 60%-80%, a video shooting frame rate of 25 frames per second, and a coverage range adjusted according to the target area.

3. The geographic information element analysis method based on a multi-modal large model according to claim 1, characterized in that: In the step S2, the remote sensing image preprocessing includes radiation correction and geometric correction. The radiation correction uses the absolute calibration method to convert the pixel gray value into the ground reflectivity based on the sensor calibration parameters. The geometric correction uses the rational polynomial coefficient model and combines with the GNSS control points for orthographic correction, and the error between the corrected image and the WGS-84 coordinate system is less than 1 pixel. The GNSS discrete point data is converted into a gridded elevation model by the inverse distance weighted interpolation method, and the grid resolution is matched with the remote sensing image.

4. The geographic feature element analysis method based on a multi-modal large model according to claim 1, characterized in that: In the step S3, a multi-modal encoder with a Transformer architecture is used to map different modal data to a 512-dimensional high-dimensional semantic space. The encoder includes an image branch, a text branch, and a numerical branch, and the outputs of each branch are fused through an attention mechanism. During feature mapping, contrastive learning is used to constrain the cross-modal semantic consistency: for different modal data of the same geographic area, the cosine similarity of their feature vectors is calculated, and the InfoNCE loss function is used to maximize the similarity of positive sample pairs and minimize the similarity of negative sample pairs.

5. The geographic feature element analysis method based on a multi-modal large model according to claim 1, characterized in that: In the step S4, first, the analysis target is converted into a task embedding vector t by a task perception module of the large model, which contains the core semantics of the geographic information elements: key elements such as land surface temperature, building density, and vegetation coverage; then, multi-modal data (infrared remote sensing image features f ir , visible light image features f vis , land surface temperature text numerical features f temp , and meteorological station wind speed features f wind ) are input into an encoder to obtain high-dimensional representations of each modality; the task perception module generates initial weights αi (0) by calculating the semantic correlation between the task embedding and the representations of each modality, for example, the initial weight of the infrared image is significantly higher than that of the visible light image because the infrared image directly reflects the land surface thermal radiation. At the same time, the system calculates the contribution degree βi of each modal to the current analysis target in real time: for infrared images, the similarity between their features and the historical heat island area infrared feature library is calculated; for ground temperature text, the deviation degree of its numerical value from the heat island threshold is calculated. For wind speed data, calculate its inhibition ability to heat diffusion; the final dynamic weight αi is determined by the initial weight αi (0) Combined with the contribution degree βi, after weight adjustment, each modal feature is weighted and fused into a comprehensive feature f according to αi fusion , which is used for subsequent heat island area identification and intensity calculation; The task-aware initial weight generation formula is: Wherein: t is the task embedding vector (dimension d, generated by the large model according to the geographic information element target, containing temperature, underlying surface type); f i is the high-dimensional feature vector of the i-th modality; the denominator is the sum of the exponents of all modalities, ensuring weight normalization; The modal contribution degree calculation formula is: β i = γ i · Sim(f i ,f i,ref ); where: γi is the modality-specific coefficient: infrared imagery γir= 1.5, land surface temperature text γtemp= 1.2, visible light imagery γvis= 0.8, embodying the prior importance of different modalities to heat island analysis); Sim(·) is the cosine similarity function, f vis is the historical heat island feature reference vector of the ith modality (e.g., the reference vector of infrared imagery is the average infrared feature of the historical heat island region). i,ref is the historical heat island feature reference vector of the ith modality (e.g., the reference vector of infrared imagery is the average infrared feature of the historical heat island region). The dynamic weight adjustment formula is: where: β is the task-aware initial weight (range 0~1) i α is the modality contribution (range 0~2, constrained by γi and similarity), adjusted weight i Dynamic amplification of the impact of high contribution modalities; The multi-modal fusion formula is: Wherein: f fusion is the integrated feature vector after fusion (dimension d), α i is the dynamically adjusted modal weight (satisfies ∑ i α i = 1), f i is the high-dimensional feature vector of each modal.

6. The geographic information element analysis method based on a multi-modal large model according to claim 1, characterized in that: In the step S5, based on the fused high-dimensional features, an improved U-Net decoder is used for element extraction. The decoder input is a 512-dimensional fused feature, and the output includes three branches: land cover classification, terrain parameter extraction, and special feature labeling. The classification layer uses a Softmax activation function, the terrain parameter layer uses a regression loss, and the labeling layer uses a Focal Loss.

7. The geographic information element analysis method based on a multi-modal large model according to claim 1, characterized in that: In the step S6, the association rule learning is based on a historical 10-year geographic event labeling dataset, which includes multiple element information such as terrain parameters, weather data, and land cover types. The rule learning module adopts an LSTM time series model, the input is a multi-element time series, and the output is the probability of event occurrence; the model training uses cross-entropy loss, the optimizer is Adam, the training rounds are 100 times, and the verification set accounts for 20%.

8. The geographic information element analysis method based on a multi-modal large model of claim 1, wherein: In step S7, verification is performed using three levels of scales: micro, meso, and macro. The micro scale is based on manual annotation labels from unmanned aerial vehicle videos, and the IoU of element extraction is calculated, with the requirement that the IoU of buildings, vegetation, and other ground objects is greater than 0.7; The meso scale uses ground measurement samples to verify the accuracy of terrain parameters and numerical elements; The macro scale compares the national basic geographic information database to calculate the overall accuracy of land cover classification.

9. The geographic information element analysis method based on a multi-modal large model of claim 1, wherein: In step S8, the analysis results are output through a visualization platform, including three types of products: hierarchical thematic maps, time series change animations, and early warning prompts. The hierarchical thematic maps are classified using the natural breakpoint method, and the color mapping uses a red-yellow-blue gradient color system. The time series change animations are generated based on the annual extraction results of the past 10 years, with a frame rate of 24 frames per second, and each frame corresponds to 1 year of data, visually displaying the process of urban expansion and vegetation change; Early warning prompts are triggered by real-time data: when new meteorological data is collected in S1, the system automatically triggers the rapid iteration of S4 to S7, and if the analysis results show that the heat island intensity exceeds the threshold, the system pushes the warning information to the user. The service output frequency is adjusted according to the data update period: regular analysis is updated daily, and early warning analysis is pushed in real time, ensuring the timeliness of the information and the value of decision support.

10. A multi-modal large model-based geographic information element analysis system, characterized in that: The system runs the geographic information element analysis method based on a multi-modal large model as claimed in any one of claims 1-9 when in use.

Citation Information

Cited By

  • Remote sensing large language model precision improvement method and device based on subdivision database and human feedback, and electronic equipment

    CN121121753A

  • Remote sensing large language model precision improvement method and device based on subdivision database and human feedback, and electronic equipment

    CN121121753B