AI-assisted natural resource factor cross-layer index association method and system

By using an AI-based method for cross-layer index association of natural resource elements, the efficiency and accuracy issues in processing remote sensing imagery and natural resource element information were addressed. A cross-layer index association model was constructed to support natural resource management and decision-making.

CN120689745BActive Publication Date: 2026-02-24GUANGZHOU DIGITAL CITIES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510695186.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-28
Publication Date
2026-02-24
Estimated Expiration
2045-05-28

AI Technical Summary

Technical Problem

Existing technologies are insufficient to efficiently and accurately process massive amounts of remote sensing imagery and complex natural resource information, and cannot fully explore the potential relationships between natural resource elements.

Method used

An AI-assisted cross-layer index association method for natural resource elements is adopted. By acquiring a set of remote sensing images and loading them into a pre-trained multimodal land cover element association model, descriptive features of remote sensing image features and natural resource element information are extracted, matching degree is calculated, layer separation and feature association are performed, and a cross-layer index association model is constructed.

Benefits of technology

It has enabled the effective correlation of natural resource elements across different data layers, providing an accurate data foundation and offering a scientific basis for natural resource management and decision-making.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689745B_ABST
    Figure CN120689745B_ABST
Patent Text Reader

Abstract

The application discloses an AI auxiliary-based natural resource factor cross-layer index correlation method and system, which comprises the following steps: firstly, a set containing multiple remote sensing images is acquired, and the set and to-be-matched natural resource factor information are loaded into a pre-trained multi-modal ground feature correlation model to respectively obtain remote sensing image features and description features; and a target remote sensing image is determined by calculating a matching degree. Different layer data is obtained by separating a target image layer, layer features are extracted, and correlation with factor index features is established, thereby constructing a cross-layer index correlation model, and effective correlation of natural resource factor cross-layer indexes is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence, and more specifically, to an AI-assisted method and system for cross-layer index association of natural resource elements. Background Technology

[0002] In the field of natural resource management, accurately linking feature indicators across different layers is crucial for the rational planning and protection of resources. Traditional methods struggle to efficiently and accurately process massive amounts of remote sensing imagery and complex natural resource information. With the development of AI technology, using AI to assist in linking natural resource indicators across different layers has become a trend. However, some existing AI-based methods have shortcomings in terms of the accuracy of feature extraction and linking, as well as model universality, failing to fully explore the potential relationships between natural resource elements. Summary of the Invention

[0003] The purpose of this invention is to provide an AI-assisted method and system for cross-layer index association of natural resource elements.

[0004] In a first aspect, embodiments of the present invention provide an AI-assisted method for cross-layer index association of natural resource elements, including:

[0005] Acquire a set of remote sensing images, wherein the set of remote sensing images includes multiple remote sensing images;

[0006] The remote sensing image set is loaded into a pre-trained multimodal ground feature association model, which is used to process the multiple remote sensing images to obtain the remote sensing image features of each remote sensing image.

[0007] The natural resource element information to be matched is loaded into the multimodal land feature association model to obtain the descriptive features corresponding to the natural resource element information. The descriptive features include baseline descriptive features, land feature features and element index features.

[0008] Calculate the matching degree between the descriptive features corresponding to the natural resource element information and the remote sensing image features of the multiple remote sensing images, and determine the target remote sensing image that matches the natural resource element information based on the matching degree;

[0009] The target remote sensing image is subjected to layer separation to obtain different types of layer data;

[0010] Extract corresponding layer features for each layer of data, and establish the correlation between each layer feature and the feature index features in the natural resource element information;

[0011] Based on the established relationships, a cross-layer indicator association model is constructed.

[0012] In a second aspect, embodiments of the present invention provide a server system, including a server, the server being used to execute the method described in the first aspect.

[0013] Compared to existing technologies, the beneficial effects provided by this invention include: Employing the AI-assisted cross-layer index association method and system for natural resource elements disclosed in this invention, a set of multiple remote sensing images is acquired, and these images are loaded with the natural resource element information to be matched into a pre-trained multimodal land cover element association model. Remote sensing image features and descriptive features are obtained separately, and the target remote sensing image is determined by calculating the matching degree. Different layer data are obtained by separating the target image layers, extracting layer features and establishing associations with element index features, thereby constructing a cross-layer index association model to achieve effective association of natural resource elements across layers. Attached Figure Description

[0014] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be considered as limiting the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 A flowchart illustrating the steps of an AI-assisted method for cross-layer index association of natural resource elements provided in an embodiment of the present invention;

[0016] Figure 2 A schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0018] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0019] In order to solve the technical problems mentioned in the background art Figure 1 This is a flowchart illustrating the AI-assisted cross-layer index association method for natural resource elements provided in this embodiment. The following is a detailed description of the AI-assisted cross-layer index association method for natural resource elements.

[0020] Step S201: Obtain a remote sensing image set, wherein the remote sensing image set includes multiple remote sensing images;

[0021] Step S202: Load the remote sensing image set into a pre-trained multimodal land cover feature association model. The multimodal land cover feature association model is used to process the multiple remote sensing images to obtain the remote sensing image features of each remote sensing image.

[0022] Step S203: Load the natural resource element information to be matched into the multimodal land feature association model to obtain the descriptive features corresponding to the natural resource element information. The descriptive features include baseline descriptive features, land feature features and element index features.

[0023] Step S204: Calculate the matching degree between the descriptive features corresponding to the natural resource element information and the remote sensing image features of the multiple remote sensing images, and determine the target remote sensing image that matches the natural resource element information based on the matching degree.

[0024] Step S205: Perform layer separation on the target remote sensing image to obtain layer data of different types;

[0025] Step S206: Extract corresponding layer features for each layer of data, and establish the correlation between each layer feature and the feature index features in the natural resource element information;

[0026] Step S207: Based on the established relationships, construct a cross-layer indicator association model.

[0027] In this embodiment of the invention, for example, the server establishes data connections with multiple satellites or multiple ground-based remote sensing devices. For instance, these satellites may be high-resolution (Gaofen) satellites specifically designed for natural resource monitoring, operating at different orbital altitudes and capable of acquiring remote sensing images of different resolutions and bands. Ground-based remote sensing devices may be distributed in specific areas, such as mountainous regions or near rivers, to supplement the satellite remote sensing images and obtain more detailed local information. The server receives remote sensing image data from these satellites and ground devices at predetermined time intervals or based on specific triggering conditions. This data is transmitted and stored in a specific format, such as the common GeoTIFF format, which includes not only the image's pixel information but also important metadata such as geographic coordinates and projection information. Over time, the server gradually accumulates a large number of remote sensing images, which together constitute a remote sensing image set. For example, during a one-month monitoring period, the server acquires remote sensing images at a 24-hour cycle, obtaining 10 4The system contains a vast collection of remote sensing images from different times and regions, covering various terrains, landforms, and natural resource distributions within the monitored area. Some images clearly show the distribution of large forests, while others highlight the course of rivers and water areas, providing a rich data foundation for subsequent analysis. After acquiring the remote sensing image set, the server loads it into a pre-trained multimodal land cover association model. Taking a specific natural resource survey project as an example, this project focuses on investigating land use types, vegetation cover, and mineral resource distribution in a specific area. This pre-trained multimodal land cover association model can perform in-depth analysis of the input remote sensing images. The model's training data comes from a large amount of historical remote sensing images and corresponding detailed natural resource element information. This data has been carefully labeled and organized, covering various land cover types and scenes. When the server loads the remote sensing image set into the model, the model begins to work. It first extracts features from each remote sensing image. For example, for a remote sensing image showing mountainous terrain, the model identifies morphological features such as mountain outlines and valley orientations. It also analyzes the image's spectral information to obtain the reflectance characteristics of different land features in various spectral bands. Through a series of complex algorithms and neural network structures, the model integrates and abstracts these features, ultimately obtaining unique remote sensing image features for each image. These features are represented as vectors, containing rich information about land features in the image, such as their category, distribution, and texture, providing a foundation for subsequent matching with natural resource element information. In actual natural resource management, specific natural resource element information often needs to be matched and analyzed with existing remote sensing images. Suppose a natural resource management department needs to understand the distribution of a specific type of mineral resource within a specific area; this generates natural resource element information to be matched. The server loads this information into a multimodal land feature association model. For example, in the case of the aforementioned mineral resources, the natural resource element information to be matched might include the type of mineral, the expected approximate distribution range, and related geological feature descriptions. After receiving this information, the model begins processing it. It extracts baseline descriptive features, land cover features, and element index features from this information through complex semantic analysis and feature extraction algorithms.Taking mineral resources as an example, the baseline descriptive features might be a general description of the mineral resource, such as "ferrous metal minerals," providing a basic semantic framework for subsequent matching. Geological features might include the geological and geomorphological features typically associated with the mineral resource, such as specific rock types and mountain range orientations. These features help in matching based on the morphology of remote sensing images. Element indicator features might consist of the mineral resource's classification code combined with specific indicator association rules, such as specific mineral codes and related reserve and grade indicator rules. These features provide a quantitative basis for matching. Through the extraction of these features, the natural resource element information to be matched is transformed into descriptive features comparable to remote sensing image features, laying the foundation for the next step of matching calculation. After obtaining the descriptive features corresponding to the natural resource element information and the remote sensing image features of each remote sensing image in the remote sensing image set, the server begins to calculate the matching degree between them. In a specific scenario, suppose the natural resource element information is a detailed description of a forest area, including the forest's tree species, approximate area, and topographic features. The remote sensing image set contains multiple images of the region and its surrounding areas. The server calculates the matching degree using specific algorithms. For example, for the semantic part of the baseline descriptive features and remote sensing image features, word vector similarity calculation methods can be used to compare the description of "forest" in the semantic space with the semantic information contained in the remote sensing image features to calculate the semantic matching degree. For land cover features, morphological algorithms can be used to calculate the morphological matching degree by comparing the expected morphology of the forest (such as shape and boundary) with the actual morphology of the land cover in the remote sensing image. For feature indicators, such as forest area, the matching degree is calculated by comparing them with the area indicators obtained from image processing and analysis of the remote sensing image. These matching degrees from different dimensions are comprehensively weighted to obtain an overall matching degree. Then, the server filters among numerous remote sensing images based on this matching degree. For example, by comparing the matching degree of all remote sensing images with the information of this natural resource feature, the remote sensing image with the highest matching degree is selected as the target remote sensing image. In this forest example, if a remote sensing image shows a high degree of matching with the forest feature information to be matched in terms of semantics, morphology, and indicators, then this image will be identified as the target remote sensing image. It is likely to accurately reflect the actual situation of the forest, providing an accurate data foundation for subsequent in-depth analysis. After identifying the target remote sensing image, the server then performs layer separation to obtain different types of layer data. Assume the target remote sensing image is a comprehensive image covering various land features such as cities, farmland, and rivers. The server first performs pixel-level semantic segmentation of the target remote sensing image based on a pre-trained land feature classification model. This pre-trained land feature classification model can label each pixel in the image with the corresponding land feature category.For example, pixels in urban areas are labeled as "urban buildings," pixels in farmland areas as "cultivated land," and pixels in river areas as "water bodies," thus generating an initial layer set containing multiple categories of land features. The server then performs spectral feature analysis and spatial topology verification on this initial layer set. Regarding spectral feature analysis, taking a specific agricultural monitoring scenario as an example, for the crop area layer, the server calculates the ratio of spectral reflectance of each pixel in the red-edge band to the near-infrared band. Because the reflectance ratio of crops with different health conditions will vary in these two bands, this method can detect the growth status of crops. Simultaneously, a dynamic threshold segmentation algorithm is used to remove abnormal spectral response areas, such as those possibly caused by sensor noise or local interference. For spatial topology verification, morphological closing operations are used to process broken patches. For example, in the urban building layer, there may be some small, broken patches due to image resolution or other reasons. Morphological closing operations can merge these small patches with the surrounding main patches, making the urban building area more complete. Furthermore, a spatial adjacency matrix is ​​constructed using Voronoi polygons. This matrix records the spatial relationships between different land cover features, such as which farmland is adjacent to rivers, and which urban areas are connected to roads. Finally, the server performs spatial overlay analysis on the validated layer data according to the land cover feature type. For example, when multiple feature types conflict at the same geographic coordinate point are detected, such as an area being labeled as both farmland and urban buildings, layer resampling is performed based on feature classification priority rules. If, in this project, urban buildings have a higher priority than farmland, the area will be redefined according to the urban building category, ultimately outputting different types of layer data with spatial topological relationships. This layer data provides a clear and accurate data foundation for subsequent layer feature extraction and correlation establishment. After acquiring different types of layer data, the server performs corresponding layer feature extraction for each layer and establishes correlations between the features of each layer and the feature indicators in the natural resource element information. Taking a land use monitoring project as an example, assuming that multiple layers of data such as cultivated land, forest land, and water area have been separated, the server performs multi-scale feature extraction for each layer. For example, for cultivated land layers, at a small scale, the focus might be on detailed features such as the boundaries and textures of individual fields; at a large scale, the focus might be on macroscopic features such as the distribution pattern of the entire cultivated land area and its relative position to surrounding features. Through this multi-scale feature extraction, the layer features corresponding to each layer of data can be comprehensively obtained. These features are represented in the form of vectors or matrices, containing rich spatial and semantic information. Next, the server calculates the cross-modal similarity between layer features and feature index features.For example, in the aforementioned land use monitoring project, feature indicators might include indicators such as cultivated land area and yield. For cultivated land area, the server quantitatively compares the regional extent features in the cultivated land layer features with the area indicator, calculating their similarity. For yield, it might combine layer features such as cultivated land soil type and irrigation conditions with the yield indicator, using a specific algorithm to calculate the similarity. Based on these cross-modal similarities, the server performs cluster analysis. It groups layer features and feature indicators with high similarity into one category, thus determining the association between each layer feature and the feature indicators in the natural resource element information. For example, if the cultivated land layer features of a certain area have high similarity to high-yield feature indicators in terms of soil fertility and irrigation water resources, then a correlation between cultivated land and high yield in that area can be established, providing strong support for subsequent analysis and decision-making. After determining the correlation between each layer feature and feature indicator feature, the server begins to construct a cross-layer indicator association model. Taking a complex natural resource integrated management scenario as an example, assume that multiple layers such as land, water resources, and vegetation have been established with corresponding feature indicators (such as land quality indicators, water flow indicators, and vegetation coverage indicators). First, the server performs multimodal feature fusion on these relationships. It integrates the relationships between features from different layers and feature indicators, generating a feature matrix containing a spatial-semantic joint representation. This matrix not only includes the spatial distribution information of each layer's features but also the semantic relationship information between them and feature indicators, such as the specific numerical relationship between land layer features and land quality indicators in a certain area, and their corresponding spatial locations. Then, cross-layer interaction modeling is performed on the feature matrix based on a multi-head attention mechanism. This mechanism captures the dynamic relationship weights between features from different layers. For example, when analyzing the relationship between water resources and vegetation, the model can dynamically adjust the relationship weights between the water resource layer and the vegetation layer according to the actual situation in different areas, thus more accurately reflecting their interaction. Next, the server adaptively aggregates the dynamic relationship weights with the feature indicators in a hierarchical manner. Based on the different feature characteristics and relationships between layers, relevant information is hierarchically integrated to generate an interpretable set of indicator association rules. For example, for an ecological environment assessment of a certain region, a rule might be generated stating that when water resource flow reaches a certain threshold and vegetation cover is within a certain range, land quality indicators will be in a good state. Subsequently, a contrastive learning strategy is used to optimize the indicator association rule set. Positive samples consist of cross-layer feature pairs associated with the same geographic entity, such as layer feature pairs of land and vegetation in the same region, which are closely related in reality. Negative samples consist of layer feature pairs with spatially conflicting distributions, such as layer feature pairs of water bodies and dry land that should not coexist in a certain region.Through this comparative learning, the indicator association rule set is continuously adjusted and optimized to make it more accurate and reliable. Finally, based on the optimized indicator association rule set, a cross-layer indicator association model incorporating a multi-layer perceptron is constructed. This model can comprehensively analyze and predict various input layer data and feature indicator characteristics, providing a scientific basis for natural resource management and decision-making. For example, in future resource planning, by inputting different land use planning schemes as layer data and combining them with existing feature indicator characteristics, the model can predict the impact of these schemes on natural resources, helping decision-makers formulate more rational plans.

[0028] In this embodiment of the invention, the multimodal feature association model is obtained in the following ways, and can be implemented through the following examples.

[0029] The sample dataset is loaded into a co-trained FLAVA model to obtain remote sensing image features, baseline description features, land cover features, and element index features of multiple training instance combinations included in the sample dataset. Each training instance combination includes a remote sensing image and natural resource element information of the remote sensing image. The multimodal land cover element association model is used to: perform land cover element parsing on the natural resource element information to obtain the land cover element set of the training instance combination; generate land cover attribute information and element index information of the land cover elements in the land cover element set of the training instance combination; and based on the training instance... The method involves obtaining the baseline descriptive features of the training instance combination based on the natural resource element information of the example combination, obtaining the land feature features of the training instance combination based on the land feature attribute information of the land feature elements in the land feature element set of the training instance combination, obtaining the element index features of the training instance combination based on the element index information of the land feature elements in the land feature element set of the training instance combination, and obtaining remote sensing image features based on the remote sensing image of the training instance combination. The land feature attribute information is used to characterize the morphological features of the land feature elements, and the element index information is composed of the land feature classification code of the land feature elements combined with index association rules.

[0030] The feature matching degree of any two training instance combinations is calculated based on the remote sensing image features, baseline description features, land cover features, and element index features of the multiple training instance combinations.

[0031] Obtain the element similarity between any two training instance combinations from the plurality of training instance combinations;

[0032] The feature consistency error is calculated based on the feature matching degree and feature similarity of any two training instances, and the multimodal feature association model is updated based on the feature consistency error.

[0033] In this embodiment of the invention, for example, the server obtains a sample dataset containing a large amount of data from different time periods and regions within a natural resource monitoring area. Each training instance combination consists of a remote sensing image and its corresponding natural resource element information. The server loads the sample dataset into a collaboratively trained FLAVA model. For example, in one training instance combination, the remote sensing image is an image of a mountainous area, and the natural resource element information describes the resources such as forest land and minerals in that area. The model performs feature analysis on the natural resource element information, identifying features such as forest land and minerals, forming a set of feature elements. For forest land, it generates its feature attribute information, such as morphological features like shape and area, as well as element indicator information, such as forest land type coding combined with indicators like forest coverage rate association rules. Baseline descriptive features, such as "distribution of natural resources in a certain mountainous area," are obtained from the overall natural resource element information. Feature features are obtained based on the feature attribute information, and element indicator features are obtained based on the element indicator information. Simultaneously, remote sensing image features are extracted from the remote sensing image. Through this processing, the server obtains various features from multiple training instance combinations. The server utilizes the features of these training instance combinations to calculate the feature matching degree between any two training instance combinations. For example, it compares a training instance combination of mountainous areas with another training instance combination containing plains and farmland. It calculates the semantic matching degree of the remote sensing image features and baseline descriptive features of both, determining the semantic similarity between the descriptions of "mountainous area" and "plains and farmland"; it calculates the morphological matching degree of the remote sensing image features and land cover features, comparing the morphological differences such as the shape of mountainous terrain and plains and farmland; and it calculates the index matching degree of the remote sensing image features and element index features, such as the matching between the forest coverage index in mountainous areas and the crop yield index in plains and farmland. Through these calculations, the server comprehensively measures the feature matching degree between the two training instance combinations. The server obtains the element similarity between any two training instance combinations using specific methods. For example, it calculates the spatial overlay analysis results of the number of land cover elements in the land cover element sets of the two training instance combinations, examining the overlap of land cover distribution; it performs element aggregation analysis to understand the spatial aggregation characteristics of land cover elements. The spatial coupling coefficient is calculated using these two analysis results to obtain the land cover element coupling degree, which is used as the element similarity. For example, in mountainous areas and plains farmland, the spatial distribution and aggregation of forest land and farmland are analyzed to derive element similarity. The server calculates element consistency error based on feature matching degree and element similarity. For instance, the errors of element similarity with semantic, morphological, and index matching degree vectors are calculated separately, and then these errors are combined to obtain the target error, which is used to calculate the element consistency error. If the element consistency error is large, it indicates that the model's current feature matching of training instance combinations deviates significantly from the actual element similarity.Based on this feature consistency error, the server updates the multimodal feature association model, adjusts the model parameters to more accurately process and match the relevant features of natural resource elements, improves model performance, and enables better association of cross-layer indicators of natural resource elements in subsequent practical applications.

[0034] In this embodiment of the invention, obtaining the element similarity between any two training instance combinations among the plurality of training instance combinations can be performed through the following example.

[0035] Based on the set of ground features combined with the multiple training instances, the coupling degree of ground features between any two training instance combinations is calculated, and the coupling degree of ground features is used as the feature similarity between any two training instance combinations.

[0036] In this embodiment of the invention, for example, it is assumed that the training instance combinations processed by the server come from monitoring data of natural resources in different regions of a province. One training instance combination, A, corresponds to remote sensing images and related natural resource element information of the northern mountainous area of ​​the province, and the set of ground features includes coniferous forests, granite veins, and other ground features. Another training instance combination, B, corresponds to the situation in the central plain of the province, and the set of ground features includes farmland, irrigation canals, and other ground features. Based on the set of ground features of these training instance combinations, the server calculates the coupling degree of ground features between any two training instance combinations, which is used as the feature similarity. Specifically, the server first performs spatial overlay analysis. For training instance combinations A and B, the server overlays the ground feature distribution layer of the mountainous area with the ground feature distribution layer of the plain. In this process, the overlap between the coniferous forests in the mountainous area and the farmland and irrigation canals in the plain is determined, and the area, shape, and other information of the overlapping area are statistically analyzed. For example, if there is a small area at the junction of the mountainous edge and the plain that is both the edge of the mountainous coniferous forest and close to the irrigation canal of the plain, this overlapping area will be recorded. Next, the server performs feature aggregation analysis. For coniferous forests in mountainous areas, it analyzes their spatial aggregation characteristics, such as whether the coniferous forests are concentrated and contiguous or scattered in small patches. For farmland in plains, it similarly analyzes their aggregation, whether they are large-scale and well-organized farmland or relatively scattered small patches. By comparing the aggregation characteristics of the two types of land features, the server understands the similarities and differences in their spatial organization. Then, based on the results of spatial overlay analysis and feature aggregation analysis, the server calculates the spatial coupling coefficient. For example, if the mountainous coniferous forests and plains farmland have little spatial overlap and their aggregation characteristics differ significantly, the calculated spatial coupling coefficient will be low, meaning the land feature coupling degree is low; conversely, if there is significant overlap at the boundary and the aggregation characteristics are somewhat similar, the spatial coupling coefficient will be high, indicating a high land feature coupling degree. Finally, the server uses the calculated land feature coupling degree as the feature similarity of training instance combinations A and B. By performing this calculation on all training instance combinations, the server can comprehensively obtain the feature similarity between any two training instance combinations, providing key data support for subsequent calculation of feature consistency error and updating the multimodal feature association model.

[0037] In this embodiment of the invention, the calculation of the coupling degree of ground features between any two training instance combinations based on the set of ground feature combinations of the plurality of training instance combinations can be performed through the following example.

[0038] The spatial overlay analysis results are used to calculate the number of ground features in the set of ground features of the first training instance combination and the set of ground features of the second training instance combination.

[0039] The feature aggregation analysis results are calculated by calculating the number of features in the feature set of the first training instance combination and the feature set of the second training instance combination.

[0040] The spatial coupling coefficients of the spatial overlay analysis results and the feature aggregation analysis results of the feature feature set of the first training instance combination and the feature feature set of the second training instance combination are calculated to obtain the feature feature coupling degree of the first training instance combination and the second training instance combination.

[0041] In this embodiment of the invention, taking the construction of a model using natural resource monitoring data of a certain region as an example, the server needs to calculate the coupling degree of ground features between any two training instance combinations based on the set of ground features from multiple training instance combinations. Assume the first training instance combination comes from the coastal area of ​​this region, and its set of ground features includes features such as beaches, reefs, and near-shore aquaculture areas; the second training instance combination comes from the inland river valley area, and its set of ground features includes features such as farmland, rivers, and villages. The server first performs a spatial overlay analysis of the number of ground features in the set of ground features of these two training instance combinations. The server precisely overlays the ground feature distribution data of the coastal area with the ground feature distribution data of the inland river valley area according to geographic coordinates. During this process, the overlap of different ground features in spatial location is statistically analyzed. For example, if a beach in the coastal area overlaps with farmland in the inland river valley area in a small area on the map, the server records the number of features for each area within this overlapping area. For the beach, the area ratio of the beach in the overlapping area may be recorded and converted into the corresponding number of features; for farmland, the number of features corresponding to its area in the overlapping area is also recorded. Through comprehensive and detailed statistics, the server obtained spatial overlay analysis results of the number of feature elements in the first and second training instance combinations, clearly showing the spatial overlap between the two sets of features. Next, the server performed feature element aggregation analysis on the feature element sets of the two training instance combinations. For beaches in coastal areas, the server analyzed their spatial aggregation characteristics, such as whether the beaches are concentrated along a long coastline or scattered into multiple small beaches; for near-shore aquaculture areas, it analyzed whether they are large-scale concentrated aquaculture or scattered small-scale aquaculture. Similarly, for cultivated land in inland river valley areas, it analyzed whether it is large, continuous cultivated land or divided into multiple small plots by rivers and roads; for villages, it analyzed whether they are concentrated large villages or scattered small village clusters. By quantifying these aggregation characteristics, such as calculating concentration and dispersion indices, the server obtained aggregation analysis data for each feature element. Combining these data yielded the feature element aggregation analysis results for the first and second training instance combinations, reflecting the spatial organization characteristics of the feature elements in the two locations. Finally, the server calculates the spatial coupling coefficient based on the spatial overlay analysis and feature aggregation analysis results obtained earlier. The calculation of the spatial coupling coefficient comprehensively considers the degree of overlap of land features in spatial location and the differences in their aggregation characteristics. For example, if the beaches in the coastal area and the cultivated land in the inland river valley area have little spatial overlap, and the concentrated distribution pattern of the beaches differs significantly from the dispersed or concentrated pattern of the cultivated land, then the calculated spatial coupling coefficient will be low, meaning that the coupling degree of land features in the combination of these two training instances is low.Conversely, if there is significant overlap in certain areas and the aggregation characteristics of ground features are somewhat similar, the spatial coupling coefficient will be high, and the coupling degree of ground feature elements will also be high. Through this calculation, the server obtains the coupling degree of ground feature elements for the first training instance combination and the second training instance combination, providing important data support for the subsequent construction of a multimodal ground feature element association model.

[0042] In this embodiment of the invention, the calculation of feature consistency error based on the feature matching degree and feature similarity of any two training instance combinations can be performed through the following example.

[0043] Generate a matching degree vector of full combination size based on the feature matching degree of any two training instances combined;

[0044] Generate a feature similarity vector of full combination size based on the feature similarity of any two training instances.

[0045] The element consistency error is calculated based on the element similarity vector and the matching degree vector.

[0046] In this embodiment of the invention, for example, assume the server is processing natural resource data from a large ecological reserve to construct a multimodal feature association model. During this process, feature consistency error needs to be calculated. The server has acquired multiple training instance combinations. For instance, training instance combination A corresponds to the wetland portion of the reserve, including wetland vegetation, water features, and related descriptions; training instance combination B corresponds to a nearby forest area, containing various trees, streams, and other features and descriptions. The server calculates the feature matching degree between any two training instance combinations, such as the semantic matching degree between remote sensing image features and baseline description features, the morphological matching degree between remote sensing image features and feature features, and the index matching degree between remote sensing image features and feature index features. After calculating the feature matching degree for all any two training instance combinations, the server generates a matching degree vector of the full combination size based on these results. Assume there are 5 different training instance combinations within the reserve, resulting in 10 possible pairwise combinations. For each combination's semantic, morphological, and index matching degree, the server organizes them into a vector form. For example, for the combination (A,B), the semantic matching degree is 0.6, the morphological matching degree is 0.5, and the index matching degree is 0.4. These values ​​are recorded sequentially at the corresponding positions in the matching degree vector. In this way, the matching degree information for all 10 combinations is recorded in the matching degree vector, which comprehensively reflects the degree of matching of various features between different training instance combinations. The server simultaneously calculates the feature similarity between any two training instance combinations, for example, by calculating the feature coupling degree of the feature sets of A and B. For the 5 training instance combinations within the protected area, there are also 10 pairwise combinations. The server organizes the feature similarity of each combination into a vector. For example, the feature similarity of combination (A,B) is calculated to be 0.35, and the server records this value at the corresponding position in the feature similarity vector. Following this method, the feature similarity records for all combinations are completed, forming a feature similarity vector of the full combination size. This vector reflects the degree of feature similarity between different training instance combinations. The server calculates the feature consistency error based on the generated feature similarity vector and matching degree vector. The server compares the semantic, morphological, and indicator matching vectors in the matching vector with the feature similarity vector, respectively. For example, it first calculates the difference between the feature similarity vector and the semantic matching vector. This is done by calculating the distribution difference of the row-level feature association distributions, comparing the distribution of each combination in terms of feature similarity and semantic matching; then, it calculates the sum of the distribution differences of all combined row set elements to obtain the first distribution difference. Similarly, it calculates the distribution difference of the column-level feature association distributions and the sum of the distribution differences of all combined row set elements to obtain the second distribution difference. The average of these two distribution differences is taken as the error between the feature similarity vector and the semantic matching vector. The errors between the feature similarity vector and the morphological matching vector, and the indicator matching vector, are calculated in the same way.Finally, these three errors are summed to obtain the target error, and the feature consistency error is calculated based on this target error. This error reflects the degree of consistency between feature matching degree and feature similarity, providing an important basis for the server to subsequently update the multimodal feature association model, so as to optimize the model and make it better associate natural resource elements.

[0047] In this embodiment of the invention, the matching degree vector includes: semantic matching degree vector, morphological matching degree vector and index matching degree vector. The calculation of the feature matching degree of any two training instance combinations based on the remote sensing image features, benchmark description features, land cover features and element index features of the multiple training instance combinations can be performed through the following example.

[0048] Calculate the semantic matching degree between the remote sensing image features and the baseline descriptive features of any two training instance combinations from the plurality of training instance combinations;

[0049] Calculate the morphological matching degree of remote sensing image features and ground feature features for any two training instance combinations from the plurality of training instance combinations;

[0050] Calculate the index matching degree of remote sensing image features and feature index features for any two training instance combinations from the plurality of training instance combinations.

[0051] The step of generating a matching degree vector of full combination size based on the feature matching degree of any two training instances includes:

[0052] Generate a semantic matching degree vector of full combination size based on the semantic matching degree of any two training instances.

[0053] A morphological matching vector of full combination size is generated based on the morphological matching degree of any two training instances combined.

[0054] Generate an index matching degree vector of full combination size based on the index matching degree of any two training instance combinations.

[0055] In this embodiment of the invention, for example, the server obtains multiple training instance combinations. For instance, training instance combination X corresponds to a city center area, including landmarks such as high-rise buildings and commercial plazas, as well as related natural resource information. Its baseline descriptive feature might be "city core commercial district." Training instance combination Y corresponds to a suburban farmland area, including landmarks such as arable land and irrigation facilities. Its baseline descriptive feature is "suburban agricultural planting area." The server calculates the semantic matching degree between the remote sensing image features and the baseline descriptive feature of any two combinations among the multiple training instance combinations. For combination (X,Y), the server compares the semantic information contained in the remote sensing image of the city center area, such as the function and distribution of buildings, with the baseline description "city core commercial district"; simultaneously, it compares the semantic information of the remote sensing image of the suburban farmland area with the baseline description "suburban agricultural planting area." Through natural language processing technology and image semantic analysis algorithms, the semantic similarity between the two is quantified to obtain the semantic matching degree. For example, the calculated semantic matching degree of combination (X,Y) is 0.2 because the city center commercial district and the suburban agricultural planting area have significant semantic differences. The server performs this calculation for all any two training instance combinations. For the same training instance combination X and Y, the server calculates the morphological matching degree between their remote sensing image features and ground feature features. In remote sensing images of urban centers, high-rise buildings exhibit tall and dense morphological features, and commercial plazas have specific shapes and layouts; while farmland in suburban areas exhibits relatively regular block-like features, and irrigation facilities have unique linear forms. The server uses image processing and pattern recognition techniques to compare the degree of conformity between the ground feature morphology in the remote sensing images of urban centers and ground feature features such as high-rise buildings and commercial plazas, and the degree of conformity between the ground feature morphology in the remote sensing images of suburban farmland and ground feature features such as farmland and irrigation facilities. For example, by calculating the similarity of morphological parameters such as shape, size, and spatial distribution, the morphological matching degree of combination (X,Y) is found to be 0.3, because the morphological differences between urban buildings and farmland are significant. Similarly, this calculation is performed for all any two training instance combinations. Next, the server calculates the index matching degree between the remote sensing image features and feature index features of any two combinations among multiple training instance combinations. Suppose that in the training instance combination X, the feature indicators of a commercial plaza might include commercial area and pedestrian traffic; while the feature indicators of suburban farmland might include arable land area and crop yield. The server extracts information related to these indicators from remote sensing imagery, such as estimating the area of ​​the commercial plaza through image analysis and estimating pedestrian traffic based on image features of vehicles and pedestrians, and compares this information with the actual feature indicators of the commercial plaza; for suburban farmland, it estimates arable land area, vegetation cover, and correlates them with crop yield indicators through imagery. The degree of matching between the two is calculated to obtain the indicator matching degree. For example, the indicator matching degree of combination (X,Y) is calculated to be 0.1, due to the significant differences between commercial and agricultural indicators.After completing the above calculations, the server generates a semantic matching degree vector of the full combination size based on the semantic matching degree of any two training instance combinations. Assuming there are 8 different combinations of training instances in the region, there are 28 possible pairwise combinations. The server records the semantic matching degree of each combination sequentially into a vector, forming a semantic matching degree vector. Similarly, a morphological matching degree vector is generated based on the morphological matching degree, and an index matching degree vector is generated based on the index matching degree. These vectors comprehensively and systematically record the degree of matching of different training instance combinations in terms of semantics, morphology, and indices, providing a crucial data foundation for subsequent calculation of feature consistency errors and optimization of the multimodal feature association model.

[0056] In this embodiment of the invention, the calculation of the element consistency error based on the element similarity vector and the matching degree vector can be performed through the following example.

[0057] Calculate the semantic consistency error between the element similarity vector and the semantic matching vector;

[0058] Calculate the morphological consistency error between the element similarity vector and the morphological matching vector;

[0059] Calculate the consistency error between the element similarity vector and the index matching vector;

[0060] The target error is obtained by summing the semantic consistency error, the morphological consistency error, and the index consistency error, and the element consistency error is calculated based on the target error.

[0061] In this embodiment of the invention, for example, the server has generated an element similarity vector and a semantic matching vector. For example, in the watershed data, training instance combination M represents mountainous forest land in the upper reaches of the river, and training instance combination N represents plain farmland in the lower reaches of the river. The element similarity vector records the element similarity of all pairwise combinations such as (M,N), and the semantic matching vector records the semantic matching degree of the corresponding combination. For combination (M,N), the element similarity is calculated through features coupling degree, etc., and is assumed to be 0.25. Regarding the semantic matching degree, the semantic matching degree of the baseline description of "forest ecosystem" in mountainous forest land and the baseline description of "agricultural planting area" in plain farmland is 0.15 after preliminary calculation. The server calculates the semantic consistency error between the element similarity vector and the semantic matching vector. First, the distribution difference of the row dimension element association distribution of the two is calculated, that is, the distribution of each combination in terms of element similarity and semantic matching degree is compared. For combination (M,N), the differences in its position and corresponding value in the element similarity vector and the semantic matching vector are analyzed. Next, calculate the sum of the distribution differences of all row set elements in the entire combination to obtain the first distribution difference. Similarly, calculate the distribution difference of the column dimension element association distribution and the sum of the distribution differences of all row set elements in the entire combination to obtain the second distribution difference. Take the average of these two distribution differences to obtain the semantic consistency error of combination (M,N). Perform this calculation for all combinations to obtain the combined semantic consistency error of the element similarity vector and semantic matching vector. Taking the same training instance combination M and N as an example, in terms of morphological matching, the remote sensing image of mountainous forest land shows dense trees and undulating terrain, matching the forest land feature characteristics; the remote sensing image of plain farmland shows large areas of regular farmland, matching the farmland feature characteristics. The previously calculated morphological matching score is 0.2. The server calculates the morphological consistency error of the element similarity vector and the morphological matching vector. Similar to the steps for calculating the semantic consistency error, compare the differences in the row and column dimension element association distributions of each combination in the element similarity vector and morphological matching vector, calculate the sum of the corresponding distribution differences, and take the average to obtain the morphological consistency error of combination (M,N). Iterate through all combinations to obtain the morphological consistency error between the feature similarity vector and the morphological matching vector. For training instance combination M, its feature indicators may include forest coverage, timber reserves, etc.; for training instance combination N, its feature indicators may include cultivated land area, grain yield, etc. Assume that the indicator matching degree of combination (M,N) calculated in the previous stage is 0.1. The server calculates the indicator consistency error between the feature similarity vector and the indicator matching vector. Still following the method of calculating semantic and morphological consistency errors, compare the distribution differences of each combination in the feature similarity vector and the indicator matching vector from the row and column dimensions, calculate the sum of the distribution differences and take the average to obtain the indicator consistency error of combination (M,N), and thus obtain the indicator consistency error between the feature similarity vector and the indicator matching vector.The server adds the semantic consistency error, morphological consistency error, and indicator consistency error to obtain the target error. For example, if the semantic consistency error is 0.3, the morphological consistency error is 0.25, and the indicator consistency error is 0.35, the target error is 0.9. Based on this target error, the feature consistency error is calculated using a specific algorithm. This feature consistency error reflects the consistency between feature similarity and matching degrees across different dimensions in the model. The server can use this to adjust and optimize the multimodal feature association model, improving the accuracy of the model's association analysis of natural resource elements.

[0062] In this embodiment of the invention, the error between the element similarity vector and the multimodal matching vector is calculated in the following manner:

[0063] Calculate the distribution difference of the row dimension element association distribution between the element similarity vector and the multimodal matching vector, wherein the multimodal matching vector is the semantic matching vector, the morphological matching vector, or the index matching vector;

[0064] The first distribution difference is obtained by summing the distribution difference of the elements in the full combination row set of the element similarity vector and the multimodal matching vector.

[0065] Calculate the distribution difference of the column dimension feature association distribution between the feature similarity vector and the multimodal matching vector;

[0066] The second distribution difference is obtained by summing the distribution differences of the elements in the full combination row set of the element similarity vector and the multimodal matching vector.

[0067] Calculate the average of the first distribution difference degree and the second distribution difference degree to obtain the error between the element similarity vector and the multimodal matching vector.

[0068] In this embodiment of the invention, for example, assume the server is processing data from a large region containing various terrain features and natural resource types, including forests, grasslands, lakes, and cities, with each region corresponding to a training instance combination. Taking the forest region (training instance combination A) and the grassland region (training instance combination B) as examples, assume the error between the feature similarity vector and the semantic matching vector is calculated first. The feature similarity vector records the pairwise feature similarity between all training instance combinations, and the semantic matching vector records the semantic matching degree between corresponding pairwise combinations. In the row dimension, for combination (A, B), the feature similarity is set to 0.3, representing the similarity between forest and grassland in terms of overall features; the semantic matching degree is 0.1, reflecting their similarity in semantic description. The server analyzes the relative position and magnitude difference of these two values ​​in their respective vector rows, quantifies this difference using a specific algorithm (such as calculating the absolute value of the difference, relative proportion, etc.), and obtains the distribution difference degree of the feature association distribution in the row dimension. This calculation is performed for all pairwise combinations (all combinations). The server sums the distribution dissimilarity scores calculated for each combination in the entire dataset along the row dimension to obtain the first distribution dissimilarity score. For example, if there are 10 training instance combinations from different regions within the entire area, there are 45 possible pairwise combinations. The server sums the distribution dissimilarity scores for each of these 45 combinations along the row dimension, assuming the total is 12.5; this is the first distribution dissimilarity score. This value reflects the overall distribution dissimilarity of the feature similarity vector and semantic matching vector along the row direction. Again, taking combination (A, B) as an example, along the column dimension, the server further analyzes the relative positions and magnitude differences of feature similarity (0.3) and semantic matching (0.1) in their respective vector columns, quantifying this difference using an algorithm similar to that used in the row dimension, to obtain the distribution dissimilarity score of the feature association distribution along the column dimension. This calculation is performed for all combinations in the entire dataset. Similar to the calculation of the first distribution dissimilarity score, the server sums the distribution dissimilarity scores calculated for each combination in the entire dataset along the column dimension to obtain the second distribution dissimilarity score. Assuming the sum of the distribution dissimilarity scores for these 45 combinations along the column dimension is 10.5, this is the second distribution dissimilarity score, reflecting the overall distribution difference between the feature similarity vector and the semantic matching vector along the column direction. The server adds the first distribution dissimilarity score of 12.5 and the second distribution dissimilarity score of 10.5, then divides by 2 to obtain an average of 11.5. This average is the error between the feature similarity vector and the semantic matching vector. In this way, the server can accurately measure the degree of difference between the feature similarity vector and the semantic matching vector. The same steps are also applied to calculating the errors between the feature similarity vector and the morphological matching vector, and between the feature similarity vector and the indicator matching vector. These errors provide important basis for the server to further optimize the multimodal feature association model, helping to improve the accuracy of the model's cross-layer indicator association analysis of natural resource elements.

[0069] In this embodiment of the invention, the calculation of the feature matching degree of any two training instance combinations based on the remote sensing image features, benchmark description features, land cover features, and element index features of the multiple training instance combinations can be performed through the following example.

[0070] The integrated descriptive features of each training instance combination are calculated based on the baseline descriptive features, land cover features, and feature index features of each training instance combination.

[0071] Calculate the feature matching degree of remote sensing image features and integrated descriptive features of any two training instance combinations from the plurality of training instance combinations;

[0072] The calculation of element consistency error based on the feature matching degree and correlation index matching degree of any two training instances includes:

[0073] A matching degree vector of full combination size is generated based on the feature matching degree of any two training instances.

[0074] Calculate the element similarity of any two training instance combinations based on the matching degree of the association index of any two training instance combinations, and generate an element similarity vector of full combination size based on the element similarity of any two training instance combinations.

[0075] The element consistency error is calculated based on the element similarity vector and the matching degree vector.

[0076] In this embodiment of the invention, for example, assume the server is processing natural resource data for a national nature reserve. This reserve includes various ecological areas such as mountains, forests, lakes, and grasslands, with each area corresponding to a training instance combination. For each training instance combination, taking the mountain area training instance combination as an example, its baseline descriptive feature might be "high-altitude mountains rich in mineral resources," and its landform features include the mountain's morphology, slope, and other geomorphic characteristics. Its element indicator features include specific indicators such as mineral types and mountain altitude. The server calculates integrated descriptive features based on these features. Through a specific algorithm, such as semantically encoding the baseline descriptive features, numerically processing the landform features and element indicator features, and then fusing these features according to certain weights. Assume the baseline descriptive feature weight is 0.3, the landform feature weight is 0.4, and the element indicator feature weight is 0.3. The semantically encoded and numerically processed features are multiplied by their respective weights and then summed to obtain the integrated descriptive features of the mountain area training instance combination. In the same way, the integrated descriptive features of all training instance combinations within the reserve are calculated. Next, the server calculates the feature matching degree between the remote sensing image features and the integrated descriptive features of any two training instance combinations from multiple training instance combinations. For example, it compares training instance combinations of mountainous areas and forest areas. The remote sensing image features of mountainous areas reflect information such as topography and vegetation cover, while the integrated descriptive features of forest areas include comprehensive information such as semantics, land features, and feature indicators. The server uses a similarity calculation algorithm to match the image with various parts of the integrated descriptive features in terms of texture, spectrum, and spatial structure. For example, by comparing the vegetation cover in the mountainous remote sensing image with the vegetation-related content in the forest integrated descriptive features, the similarity between the two is calculated to obtain the feature matching degree of these two training instance combinations. This calculation is performed for all any two training instance combinations. Based on the feature matching degree of any two training instance combinations, the server generates a matching degree vector of the full combination size. Assuming there are 8 different ecological regions within the protected area, there are 28 possible pairwise combinations. The server records the feature matching degree of each combination sequentially into a vector, forming a matching degree vector that comprehensively records the feature matching between different combinations. The server calculates feature similarity based on the correlation indicator matching degree of any two training instance combinations. For example, there is a certain correlation between mineral indicators in mountainous areas and timber resource indicators in forestous areas. By analyzing the matching of these correlated indicators and combining methods such as the coupling degree of land features, the element similarity is calculated. Assuming the element similarity of the mountain and forest combination is 0.25, after calculating the element similarity for all combinations, a feature similarity vector of the full combination size is generated. Again, taking 8 training instance combinations as an example, the element similarity of 28 combinations is recorded sequentially in the vector. Finally, the server calculates the element consistency error based on the element similarity vector and the matching degree vector.The method described earlier for calculating the distribution differences in row and column dimensions and then averaging them is used to calculate the error between the two. For example, the distribution difference of the row dimension feature association distribution is calculated first, and then the sum of the distribution differences of all combined row set elements is calculated to obtain the first distribution difference; similarly, the column dimension correlation value is calculated to obtain the second distribution difference. The average of the two is then used to obtain the feature consistency error. This error reflects the consistency between feature matching degree and feature similarity in the model. The server can use this to optimize the multimodal feature association model and improve the accuracy of cross-layer index association analysis of natural resource elements within nature reserves.

[0077] In this embodiment of the invention, the calculation of the integrated descriptive features of each training instance combination based on the baseline descriptive features, land cover features, and feature index features of each training instance combination can be performed through the following example.

[0078] The integrated descriptive features of each training instance combination are obtained by performing multimodal feature balancing aggregation or multimodal feature weight aggregation on the baseline descriptive features, land cover features, and feature index features of each training instance combination.

[0079] In this embodiment of the invention, for example, it is assumed that the server processes large-scale regional data covering various geographical environments and natural resource types. This region includes different areas such as mountains, plains, rivers, and wetlands, with each area constituting a training instance combination. Taking the mountain training instance combination as an example, its baseline descriptive feature is "high-altitude mountainous area rich in metallic minerals," and its landform features include landform-related characteristics such as mountain morphology and rock texture. Its element indicator features include specific numerical indicators such as mineral reserves and mountain altitude. During multimodal element balancing aggregation, the server assigns equal weights to these three features. First, the baseline descriptive feature is processed into text vectors, converting textual information into numerical vectors. For example, word embedding technology is used to convert "high-altitude mountainous area rich in metallic minerals" into a vector of a specific dimension. For landform features, image processing technology is used to quantify information such as mountain morphology and rock texture into numerical vectors. The element indicator features are themselves in numerical form and can be used directly. Then, the server performs a simple averaging operation on these three numerical vectors. For example, assuming the baseline descriptive feature vector is [0.2, 0.3, 0.4], the land feature vector is [0.1, 0.4, 0.3], and the element indicator feature vector is [0.3, 0.2, 0.5], the integrated descriptive feature vector obtained through balanced aggregation is [(0.2+0.1+0.3) / 3, (0.3+0.4+0.2) / 3, (0.4+0.3+0.5) / 3], i.e., [0.2, 0.3, 0.4]. The server performs multimodal element balanced aggregation on all training instance combinations within the region in the same manner to obtain the integrated descriptive features for each combination. Taking the mountainous training instance combination as an example again, considering that the region mainly focuses on mineral resource development, the server decides to assign a higher weight of 0.5 to the element indicator features, a weight of 0.3 to the baseline descriptive features, and a weight of 0.2 to the land feature features. Similarly, the baseline descriptive features are first vectorized from text, the land feature features are numericalized, and the element indicator features are kept in numerical form. Assume the baseline descriptive feature vector is [0.2, 0.3, 0.4], the land feature vector is [0.1, 0.4, 0.3], and the element index feature vector is [0.3, 0.2, 0.5]. Through weight aggregation, the integrated descriptive feature vector is calculated as [(0.2×0.3+0.1×0.2+0.3×0.5), (0.3×0.3+0.4×0.2+0.2×0.5), (0.4×0.3+0.3×0.2+0.5×0.5)], which is [0.23, 0.27, 0.43]. The server flexibly adjusts the weights based on the characteristics and analytical needs of different training instance combinations, performing multimodal element weight aggregation on all training instance combinations within the region to obtain the integrated descriptive features for each combination. These integrated descriptive features integrate multiple aspects of information, laying the foundation for subsequent calculations of feature matching degrees between training instance combinations and optimization of the multimodal land feature association model.

[0080] In this embodiment of the invention, before updating the multimodal feature association model based on the feature consistency error, the following implementation method is also provided.

[0081] Calculate the standard feature matching degree of any two training instance combinations based on the standard feature labels of the multiple training instance combinations;

[0082] The comparison error is calculated based on the feature matching degree and standard element matching degree of any two training instance combinations.

[0083] The step of updating the multimodal feature association model based on the feature consistency error includes:

[0084] The final error of the multimodal feature association model is calculated based on the feature consistency error and the comparison error.

[0085] The multimodal feature association model is updated based on the final error.

[0086] In this embodiment of the invention, for example, assume that the server is processing natural resource data of a large ecological park. This park includes various ecological areas such as forests, wetlands, and farmland, each corresponding to a training instance combination used to construct a multimodal feature association model. The server obtains standard feature labels for multiple training instance combinations. For example, the standard feature labels for the forest area training instance combination specify information such as the tree species types and coverage standard range of the forest; the standard feature labels for the wetland area training instance combination specify the water quality standards and biodiversity indicators of the wetland. The server calculates the standard feature matching degree between any two training instance combinations based on these standard feature labels. Taking the forest and wetland training instance combinations as an example, the server compares the tree species types of the forest with the suitable tree species for growth in the wetland (if there are relevant standard associations), as well as the forest coverage standard range and the impact standards of the wetland ecology on the surrounding vegetation cover. Through a specific algorithm, the matching degree of the two in the standard feature labels is quantified to obtain the standard feature matching degree. For example, if the calculated standard feature matching degree of the forest and wetland training instance combination is 0.2, it indicates that the two have a low degree of similarity in terms of standard features. The server performs this calculation for all any two training instance combinations. The server has already calculated the feature matching degree for any two training instance combinations. For example, the feature matching degree for the forest and wetland training instance combination, based on remote sensing image features, baseline description features, and other land cover features, has been calculated to be 0.3. The server calculates the comparison error based on the feature matching degree and the standard feature matching degree. By analyzing the difference between the feature matching degree and the standard feature matching degree, a specialized error calculation method is used, such as calculating the sum of the squares of the differences and then averaging them. Assume that the calculated comparison error for the forest and wetland training instance combination is 0.05. This calculation is performed for all any two training instance combinations to obtain the overall comparison error. The server has obtained the feature consistency error. For example, by calculating the error between the feature similarity vector and the semantic, morphological, and index matching degree vectors, the overall feature consistency error is calculated to be 0.1. The server calculates the final error of the multimodal land cover feature association model based on the feature consistency error and the comparison error. Assuming a simple summation method is used (in practice, a more complex calculation method may be employed depending on the model's characteristics), the average comparison error (assumed to be 0.08) obtained by summing the feature consistency error (0.1) and the comparison error is added together, resulting in a final error of 0.18. The server updates the multimodal feature association model based on this final error. The model contains a series of parameters, such as weights in a neural network. The server uses optimization algorithms, such as stochastic gradient descent, to adjust these parameters based on the final error.This enables the model to reduce the difference between feature matching degree and standard feature matching degree, as well as the inconsistency between feature similarity and each matching degree vector, when processing similar training instance combinations in the future. This improves the accuracy and reliability of the model's cross-layer index correlation analysis of natural resource elements in ecological parks, and better supports the management and planning of ecological parks.

[0087] In this embodiment of the invention, the calculation of the comparison error based on the feature matching degree and standard element matching degree of any two training instance combinations can be performed through the following example.

[0088] Generate a matching degree vector of full combination size based on the feature matching degree of any two training instances combined;

[0089] Based on the standard feature matching degree of any two training instances, a standard feature identifier vector of full combination size is generated. In the standard feature identifier vector, the standard feature matching degree on the self-association diagonal of the feature is set as a valid identifier, and the other standard feature matching degrees are reset to invalid identifiers.

[0090] The comparison error is calculated based on the matching degree vector and the standard element identifier vector.

[0091] In this embodiment of the invention, for example, suppose the server is processing data from a large nature reserve, which includes various ecological regions such as grasslands, forests, and lakes, with each ecological region corresponding to a training instance combination. The server has calculated the feature matching degree between any two training instance combinations. For example, for the grassland and forest training instance combination, by comparing their remote sensing image features, baseline description features, land cover features, and element index features, the feature matching degree is found to be 0.35; the feature matching degree for the grassland and lake training instance combination is 0.2, and so on. The server generates a matching degree vector of the full combination size based on the feature matching degree of all any two training instance combinations. Suppose there are 5 different ecological region training instance combinations within the reserve, resulting in 10 possible pairwise combinations. The server records the feature matching degree of these 10 combinations sequentially into a vector, forming a matching degree vector. For example, the matching degree vector might be [0.35, 0.2, 0.4, 0.15, 0.3, 0.25, 0.1, 0.45, 0.28, 0.32], which comprehensively records the feature matching degree between different training instance combinations. Each training instance combination has a corresponding standard feature label. For example, the standard feature labels for a grassland training instance combination include grass species type and vegetation coverage standard; the standard feature labels for a forest training instance combination include tree species type and canopy closure standard. The server generates a standard feature identifier vector of full combination size based on the standard feature matching degree of any two training instance combinations. During the generation process, the standard feature matching degree on the self-associative diagonal (i.e., the matching degree between the same training instance combination and itself) is set as a valid identifier. For example, the standard feature matching degree between grassland and itself is set to 1 (representing a valid identifier), and the standard feature matching degree between forest and itself is also set to 1. However, for other non-self-associative standard feature matching degrees, such as the standard feature matching degrees of combinations of grassland and forest, or grassland and lake, they are reset to invalid identifiers and set to 0. Assuming that the standard feature identifier vector generated by 5 training instance combinations is [1,0,0,0,0,0,1,0,0,0,0,0,1,0,0,0,0,0,1,0,0,0,0,0,1], this vector highlights the self-associative standard feature matching situation. The server calculates the comparison error based on the matching degree vector and the standard feature identifier vector. The server uses a specific algorithm, such as calculating the sum of squares of the differences between corresponding elements of the two vectors. The matching degree vector [0.35,0.2,0.4,0.15,0.3,0.25,0.1,0.45,0.28,0.32] is subtracted from the corresponding elements of the standard feature identifier vector [1,0,0,0,0,0,1,0,0,0,0,0,1,0,0,0,0,0,1,0,0,0,0,0,1], and the sum is then calculated. The resulting value is the comparison error.This comparison error reflects the degree of difference between the feature matching degree and the standard feature matching degree, providing an important basis for the server to further update the multimodal land feature association model, so as to improve the accuracy of the model's cross-layer index association analysis of natural resource elements in nature reserves.

[0092] In this embodiment of the invention, the step of separating the target remote sensing image into layers to obtain different types of layer data can be implemented through the following examples.

[0093] The target remote sensing image is segmented into pixels at the semantic level based on a pre-trained land feature classification model to generate an initial layer set containing multiple categories of land features.

[0094] The initial layer set is subjected to spectral feature analysis and spatial topology verification. The spectral feature analysis includes calculating the ratio of spectral reflectance of each pixel in the red-edge band to the near-infrared band, and removing abnormal spectral response regions based on a dynamic threshold segmentation algorithm. The spatial topology verification includes processing broken patches using morphological closing operations and constructing a spatial adjacency matrix between features using Voronoi polygons.

[0095] The verified layer data is spatially overlaid according to the type of geographic feature. When a conflict of multiple feature types is detected at the same geographic coordinate point, the layer is resampled based on the feature classification priority rule, and the layer data of the different types with spatial topological relationships are output.

[0096] In this embodiment of the invention, for example, suppose the server is processing a remote sensing image of a target area, which encompasses various land cover types such as cities, villages, farmland, forests, and rivers. The aim is to obtain different types of layer data by separating the target remote sensing image into layers for subsequent analysis of natural resource elements. The server uses a pre-trained land cover classification model to perform pixel-level semantic segmentation on the target remote sensing image. This land cover classification model has been trained on a large amount of accurately labeled remote sensing image data and can accurately identify various land cover elements. For example, in this target remote sensing image, the model identifies pixels in urban areas as "buildings," pixels in farmland areas as "cultivated land," pixels in forest areas as "vegetation," and pixels in river areas as "water bodies," etc. By classifying each pixel in the image, an initial layer set containing multiple categories of land cover elements is generated. In this set, each layer corresponds to a land cover category, such as "building layer," "cultivated land layer," "vegetation layer," "water body layer," etc., and each layer records the pixel distribution of the corresponding land cover element in the image. The server performs spectral feature analysis on the initial layer set. Taking the "vegetation layer" as an example, the spectral feature analysis requires calculating the ratio of spectral reflectance of each pixel in the red-edge band to the near-infrared band. This is because the reflectance ratios of different vegetation types and health conditions vary between these two bands. For example, healthy green vegetation has higher reflectance in the near-infrared band and relatively lower reflectance in the red-edge band; calculating this ratio effectively distinguishes different vegetation conditions. Simultaneously, the server uses a dynamic threshold segmentation algorithm to remove abnormal spectral response regions. During image acquisition, some pixels may exhibit abnormal spectral responses due to sensor noise, atmospheric interference, or local special conditions. The dynamic threshold segmentation algorithm dynamically determines a threshold based on the overall spectral characteristics of the image. Pixel regions with spectral reflectance ratios exceeding this threshold are identified as abnormal spectral response regions and removed. For example, in the "vegetation layer," some pixels may have spectral reflectance ratios that significantly deviate from the normal vegetation range, possibly due to local shadows or sensor malfunctions. These regions will be identified and removed by the algorithm, thereby improving the quality of the layer data. The server also performs spatial topology verification on the initial layer set. On one hand, morphological closing operations are used to process fragmented patches. In the "Building Layer," due to image resolution limitations or other factors, there may be some isolated and fragmented small patches. These small patches may be local details of buildings that have been segmented due to resolution issues, rather than independent features. Morphological closing operations, through dilation followed by erosion, merge these fragmented small patches with the surrounding main patches, making the building boundaries more continuous and complete, conforming to the actual spatial morphology. On the other hand, the server constructs a spatial adjacency matrix between features using Voronoi polygons.Taking the "farmland layer" and "road layer" as examples, the Voronoi polygon algorithm divides the space into multiple polygonal regions based on the location of each feature. Points within each polygon are closest to their corresponding features. This clearly defines the spatial adjacency relationships between different features. For example, it identifies which farmlands are adjacent to roads, which forests border rivers, etc., and records these relationships in a matrix. This spatial adjacency matrix provides crucial information for subsequent analysis of the interactions and spatial layout of features. The server performs spatial overlay analysis on the validated layer data according to feature type. During the analysis, conflicts between multiple feature types at the same geographic coordinate point may be detected. For example, in a certain area, the "building layer" and "farmland layer" overlap at some coordinate points, meaning the area is identified as both buildings and farmland. In this case, the server resamples the layers based on feature classification priority rules. Assume that in the feature classification priority rules for this project, buildings have a higher priority than farmland. For conflicting coordinate points, the server will redefine the area according to building type, classifying it into the "building layer" and discarding the farmland identifier. This process eliminates conflicts between feature types. Ultimately, the server outputs layer data of different types with spatial topological relationships. This layer data not only accurately reflects the distribution of different land features but also contains topological information such as their spatial adjacency relationships, laying a solid data foundation for subsequent feature extraction, establishing correlations with natural resource element information, and constructing cross-layer indicator correlation models.

[0097] In this embodiment of the invention, the step of extracting corresponding layer features for each layer of data and establishing the correlation between each layer feature and the feature index features in the natural resource element information can be implemented through the following example.

[0098] Perform multi-scale feature extraction on each layer of data to obtain the layer features corresponding to each layer of data;

[0099] Calculate the cross-modal similarity between the layer features and the feature index features;

[0100] Cluster analysis is performed based on the cross-modal similarity to determine the correlation between the features of each layer and the feature indicators in the natural resource element information.

[0101] In this embodiment of the invention, for example, assume that the server is processing remote sensing image data of a large agricultural area, which includes various land features such as arable land, irrigation facilities, and roads, and the corresponding natural resource element information is known, including characteristic indicators such as crop yield and irrigation water consumption. The server performs multi-scale feature extraction for each layer of data. Taking the arable land layer as an example, at a small scale, the server focuses on the detailed features of individual fields. For example, using high-resolution image information, it extracts the clarity of field boundaries and the texture features of the soil within the field. These small-scale features can reflect the refined management of farmland; for example, the regularity of field boundaries may indicate the rationality of farmland planning. At a large scale, the server focuses on the macroscopic features of the entire arable land area. For example, it analyzes the overall distribution pattern of arable land, whether it is concentrated and contiguous or scattered, and the relative positional relationship between arable land and surrounding land features (such as irrigation facilities and roads). Large-scale features help to grasp the spatial layout of arable land within the entire area and its interrelationship with other elements. By comprehensively extracting features from different scales, the server obtains the layer features corresponding to the cultivated land layer. These features are represented in vector or matrix form, comprehensively covering various information about cultivated land from micro to macro levels. Similarly, the server performs similar multi-scale feature extraction for irrigation facility layers and road layers to obtain their respective layer features. The server then calculates the cross-modal similarity between layer features and feature index features. For cultivated land layer features and the feature index feature of crop yield, the server employs a specific algorithm. For example, considering that layer features such as cultivated land area and soil quality are closely related to crop yield, the server quantifies the features in the cultivated land layer that reflect this information and compares them with the crop yield index. Suppose that analysis reveals a functional relationship between the soil fertility feature value in the cultivated land layer of a certain region and the crop yield of that region, the server uses this relationship to calculate the similarity between the two. For irrigation facility layer features and irrigation water consumption feature index features, the server analyzes the correlation between layer features such as the coverage and water delivery capacity of irrigation facilities and irrigation water consumption, and derives the cross-modal similarity by comparing the numerical and logical connections between the two. Similarly, cross-modal similarity is calculated for all layer features and their corresponding element index features. Based on the calculated cross-modal similarity, the server performs cluster analysis. For example, the server organizes the cross-modal similarity data of all cultivated land layer features and crop yield. If it finds that the cultivated land layer features of certain areas have a high similarity to the element index features of high-yield crops in terms of soil fertility, irrigation conditions, etc., the server groups these areas with similar features into one category. Similarly, for the layer features of irrigation facilities and the element index features of irrigation water consumption, areas with high similarity to different water consumption demands in terms of facility coverage, water conveyance efficiency, etc., are clustered separately.Through such cluster analysis, the server can clearly determine the correlation between the features of each layer and the feature indicators in the natural resource element information. For example, it clarifies which specific layer features of cultivated land correspond to high yields, and which irrigation facility features are associated with high irrigation water consumption, providing a crucial basis for subsequently constructing cross-layer indicator correlation models.

[0102] In this embodiment of the invention, the construction of a cross-layer indicator association model based on the established association relationship can be implemented through the following example.

[0103] Multimodal feature fusion is performed on the correlation between the features of each layer and the feature index features to generate a correlation feature matrix containing spatial-semantic joint representation;

[0104] The cross-layer interaction modeling of the associated feature matrix is ​​performed based on the multi-head attention mechanism to capture the dynamic association weights between features of different layers.

[0105] The dynamic association weights and feature indicators are adaptively aggregated in a hierarchical manner to generate an interpretable set of indicator association rules.

[0106] The indicator association rule set is optimized by a contrastive learning strategy, wherein positive samples consist of cross-layer feature pairs associated with the same geographic entity, and negative samples consist of layer feature pairs with spatially conflicting distributions.

[0107] A cross-layer indicator association model incorporating a multi-layer perceptron is constructed based on the optimized indicator association rule set.

[0108] In this embodiment of the invention, for example, assume the server is processing data from a large ecological reserve. This reserve contains various land features such as forests, rivers, and wetlands, corresponding to multiple data layers, and various natural resource element indicators, such as forest coverage and river flow, are known. The server first performs multimodal feature fusion on the correlation between the features of each layer and the element indicator features. For example, forest layer features include information such as tree species and distribution density, river layer features include river width and flow velocity, while forest coverage and river flow are the corresponding element indicator features. The server integrates these different types of features. For the forest layer, tree species are represented in semantic vector form, and distribution density is presented in numerical form; the width and flow velocity of the river layer are also converted into numerical values. At the same time, element indicator features such as forest coverage and river flow are also incorporated. Through a specific algorithm, these features from different modalities (semantic, numerical, etc.) are combined to generate a correlation feature matrix containing a spatial-semantic joint representation. Each row of the matrix may represent a specific region, and the columns correspond to different features, reflecting not only the spatial distribution of each layer feature but also integrating semantic information and the correlation with element indicator features. The server uses a multi-head attention mechanism to perform cross-layer interactive modeling of the associated feature matrix. Taking the forest and river layers as an example, the multi-head attention mechanism allows the model to focus on the relationships between features in the forest and river layers from different perspectives. In a certain area, the presence of forests may affect the water quality and flow of rivers, while rivers provide water for forests. Through the multi-head attention mechanism, the model analyzes the interactions between forest vegetation type, area, and river flow and velocity from multiple dimensions. In this process, the dynamic association weights between features of different layers are captured. For example, in arid areas, the influence of forest vegetation on river flow may have a higher weight; while in humid areas, the weight of river water supply to forests may be more significant. In this way, the model can dynamically determine the weights of important associations between layers in different regions. The server adaptively aggregates the dynamic association weights with feature indicators in a hierarchical manner. For example, based on the dynamic association weights between forests and rivers in different regions, as well as feature indicators such as forest cover and river flow, an interpretable set of indicator association rules is generated. For areas with high forest cover and close to rivers, the rule might be "high forest cover combined with appropriate river flow helps maintain good biodiversity." This rule is derived from the interrelationships between features and element indicators across different layers, through hierarchical aggregation using dynamic association weights. It clearly explains the connections and influences between different natural resource elements. The server optimizes the indicator association rule set through a contrastive learning strategy. Positive samples consist of cross-layer feature pairs associated with the same geographic entity, such as a forest and its surrounding rivers that are affected by it; these are ecologically interconnected and constitute positive samples.Negative samples consist of layer feature pairs with conflicting spatial distributions, such as forests and large areas of exposed rock that are impossible to coexist in reality (assuming such a situation does not exist in the protected area). The server uses these positive and negative samples to optimize the indicator association rule set. For example, if a rule in the rule set indicates that forest cover and river flow are not correlated, but positive samples show a clear correlation, the server will adjust that rule. By continuously comparing the differences between positive and negative samples and the rule set, the rule set becomes more accurate and complete. Based on the optimized indicator association rule set, the server constructs a cross-layer indicator association model incorporating a multilayer sensing network. The multilayer sensing network can learn and apply complex indicator association rules. For example, when new remote sensing image data of a certain area of ​​the protected area is input, after layer separation and feature extraction, the model, based on the optimized indicator association rule set and through calculation by the multilayer sensing network, can predict the characteristics of natural resource elements in that area, such as predicting the impact of changes in forest cover on river ecology, or inferring the ecological status of surrounding wetlands based on changes in river flow. This model provides a powerful analytical tool for natural resource management and ecological research in protected areas.

[0109] This invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned AI-assisted cross-layer index association method for natural resource elements. Figure 2 As shown, Figure 2 This is a structural block diagram of a computer device 100 provided in an embodiment of the present invention. The computer device 100 includes a memory 111, a processor 112, and a communication unit 113.

[0110] To enable data transmission or interaction, the memory 111, processor 112, and communication unit 113 are electrically connected to each other, either directly or indirectly. For example, these components can be electrically connected to each other via one or more communication buses or signal lines.

[0111] For illustrative purposes, the foregoing description has been made with reference to specific embodiments. However, the foregoing illustrative discussions are not intended to be exhaustive or to limit the present disclosure to the precise forms disclosed. Numerous modifications and variations are possible in accordance with the foregoing teachings. These embodiments were chosen and described in order to best illustrate the principles of the present disclosure and its practical application, thereby enabling those skilled in the art to best utilize the disclosure and to employ various embodiments with different modifications to suit a particular intended application.

Claims

1. An AI-assisted method for cross-layer index association of natural resource elements, characterized in that, include: Acquire a set of remote sensing images, wherein the set of remote sensing images includes multiple remote sensing images; The remote sensing image set is loaded into a pre-trained multimodal ground feature association model, which is used to process the multiple remote sensing images to obtain the remote sensing image features of each remote sensing image. The natural resource element information to be matched is loaded into the multimodal land feature association model to obtain the descriptive features corresponding to the natural resource element information. The descriptive features include baseline descriptive features, land feature features, and element index features. Calculate the matching degree between the descriptive features corresponding to the natural resource element information and the remote sensing image features of the multiple remote sensing images, and determine the target remote sensing image that matches the natural resource element information based on the matching degree; The target remote sensing image is subjected to layer separation to obtain different types of layer data; Extract corresponding layer features for each layer of data, and establish the correlation between each layer feature and the feature index features in the natural resource element information; Based on the established relationships, construct a cross-layer indicator association model; The step of extracting corresponding layer features for each layer of data and establishing the correlation between each layer feature and the feature index features in the natural resource element information includes: Perform multi-scale feature extraction on each layer of data to obtain the layer features corresponding to each layer of data; Calculate the cross-modal similarity between the layer features and the feature index features; Cluster analysis is performed based on the cross-modal similarity to determine the correlation between the features of each layer and the feature indicators in the natural resource element information; The step of constructing a cross-layer indicator association model based on the established relationships includes: Multimodal feature fusion is performed on the correlation between the features of each layer and the feature index features to generate a correlation feature matrix containing spatial-semantic joint representation; The cross-layer interaction modeling of the associated feature matrix is ​​performed based on the multi-head attention mechanism to capture the dynamic association weights between features of different layers. The dynamic association weights and feature indicators are adaptively aggregated in a hierarchical manner to generate an interpretable set of indicator association rules. The indicator association rule set is optimized by a contrastive learning strategy, wherein positive samples consist of cross-layer feature pairs associated with the same geographic entity, and negative samples consist of layer feature pairs with spatially conflicting distributions. A cross-layer indicator association model incorporating a multi-layer perceptron is constructed based on the optimized indicator association rule set.

2. The method according to claim 1, characterized in that, The multimodal feature association model is obtained through the following methods: The sample dataset is loaded into a co-trained FLAVA model to obtain remote sensing image features, baseline description features, land cover features, and element index features of multiple training instance combinations included in the sample dataset. Each training instance combination includes a remote sensing image and natural resource element information of the remote sensing image. The multimodal land cover element association model is used to: perform land cover element parsing on the natural resource element information to obtain the land cover element set of the training instance combination; generate land cover attribute information and element index information of the land cover elements in the land cover element set of the training instance combination; and based on the training instance... The method involves obtaining the baseline descriptive features of the training instance combination based on the natural resource element information of the example combination, obtaining the land feature features of the training instance combination based on the land feature attribute information of the land feature elements in the land feature element set of the training instance combination, obtaining the element index features of the training instance combination based on the element index information of the land feature elements in the land feature element set of the training instance combination, and obtaining remote sensing image features based on the remote sensing image of the training instance combination. The land feature attribute information is used to characterize the morphological features of the land feature elements, and the element index information is composed of the land feature classification code of the land feature elements combined with index association rules. The feature matching degree of any two training instance combinations is calculated based on the remote sensing image features, baseline description features, land cover features, and element index features of the multiple training instance combinations. The spatial overlay analysis results are used to calculate the number of ground features in the set of ground features of the first training instance combination and the set of ground features of the second training instance combination. The feature aggregation analysis results are calculated by calculating the number of features in the feature set of the first training instance combination and the feature set of the second training instance combination. Calculate the spatial coupling coefficient of the spatial overlay analysis results and the feature aggregation analysis results of the feature feature set of the first training instance combination and the feature feature set of the second training instance combination, and obtain the feature feature coupling degree of the first training instance combination and the second training instance combination. Use the feature feature coupling degree as the feature similarity of any two training instance combinations. Generate a matching degree vector of full combination size based on the feature matching degree of any two training instances combined; Generate a feature similarity vector of full combination size based on the feature similarity of any two training instances. The feature consistency error is calculated based on the feature similarity vector and the matching degree vector, and the multimodal feature association model is updated based on the feature consistency error.

3. The method according to claim 2, characterized in that, The matching degree vector includes: semantic matching degree vector, morphological matching degree vector, and index matching degree vector. The calculation of the feature matching degree between any two training instance combinations based on the remote sensing image features, baseline descriptive features, land cover features, and element index features of the multiple training instance combinations includes: Calculate the semantic matching degree between the remote sensing image features and the baseline descriptive features of any two training instance combinations from the plurality of training instance combinations; Calculate the morphological matching degree of remote sensing image features and ground feature features for any two training instance combinations from the plurality of training instance combinations; Calculate the index matching degree of remote sensing image features and feature index features for any two training instance combinations from the plurality of training instance combinations. The step of generating a matching degree vector of full combination size based on the feature matching degree of any two training instances includes: Generate a semantic matching degree vector of full combination size based on the semantic matching degree of any two training instances. A morphological matching vector of full combination size is generated based on the morphological matching degree of any two training instances combined. Generate an index matching degree vector of full combination size based on the index matching degree of any two training instance combinations.

4. The method according to claim 3, characterized in that, The calculation of feature consistency error based on the feature similarity vector and the matching degree vector includes: Calculate the semantic consistency error between the element similarity vector and the semantic matching vector; Calculate the morphological consistency error between the element similarity vector and the morphological matching vector; Calculate the consistency error between the element similarity vector and the index matching vector; The target error is obtained by summing the semantic consistency error, the morphological consistency error, and the index consistency error, and the element consistency error is calculated based on the target error.

5. The method according to claim 4, characterized in that, The error between the feature similarity vector and the multimodal matching vector is calculated as follows: Calculate the distribution difference of the row dimension element association distribution between the element similarity vector and the multimodal matching vector, wherein the multimodal matching vector is the semantic matching vector, the morphological matching vector, or the index matching vector; The first distribution difference is obtained by summing the distribution difference of the elements in the full combination row set of the element similarity vector and the multimodal matching vector. Calculate the distribution difference of the column dimension feature association distribution between the feature similarity vector and the multimodal matching vector; The second distribution difference is obtained by summing the distribution differences of the elements in the full combination row set of the element similarity vector and the multimodal matching vector. Calculate the average of the first distribution difference degree and the second distribution difference degree to obtain the error between the element similarity vector and the multimodal matching vector.

6. The method according to claim 2, characterized in that, The calculation of the feature matching degree between any two training instance combinations based on the remote sensing image features, baseline description features, land cover features, and element index features of the multiple training instance combinations includes: The integrated descriptive features of each training instance combination are obtained by performing multimodal feature balancing aggregation or multimodal feature weight aggregation on the baseline descriptive features, land cover features, and feature index features of each training instance combination. Calculate the feature matching degree of remote sensing image features and integrated descriptive features of any two training instance combinations from the plurality of training instance combinations; The calculation of element consistency error based on the feature matching degree and association index matching degree of any two training instance combinations includes: A matching degree vector of full combination size is generated based on the feature matching degree of any two training instances combined. Calculate the element similarity of any two training instance combinations based on the matching degree of the association index of any two training instance combinations, and generate an element similarity vector of full combination size based on the element similarity of any two training instance combinations. The element consistency error is calculated based on the element similarity vector and the matching degree vector.

7. The method according to claim 2, characterized in that, Before updating the multimodal feature association model based on the feature consistency error, the method further includes: Calculate the standard feature matching degree of any two training instance combinations based on the standard feature labels of the multiple training instance combinations; Generate a matching degree vector of full combination size based on the feature matching degree of any two training instances combined; Based on the standard feature matching degree of any two training instances, a standard feature identifier vector of full combination size is generated. In the standard feature identifier vector, the standard feature matching degree on the self-association diagonal of the feature is set as a valid identifier, and the other standard feature matching degrees are reset to invalid identifiers. The comparison error is calculated based on the matching degree vector and the standard element identifier vector. The step of updating the multimodal feature association model based on the feature consistency error includes: The final error of the multimodal feature association model is calculated based on the feature consistency error and the comparison error. The multimodal feature association model is updated based on the final error.

8. The method according to claim 1, characterized in that, The process of separating layers in the target remote sensing image to obtain different types of layer data includes: The target remote sensing image is segmented into pixels at the semantic level based on a pre-trained ground feature classification model to generate an initial layer set containing multiple categories of ground features. The initial layer set is subjected to spectral feature analysis and spatial topology verification. The spectral feature analysis includes calculating the ratio of spectral reflectance of each pixel in the red-edge band to the near-infrared band, and removing abnormal spectral response regions based on a dynamic threshold segmentation algorithm. The spatial topology verification includes processing broken patches using morphological closing operations and constructing a spatial adjacency matrix between features using Voronoi polygons. The verified layer data is spatially overlaid according to the type of geographic feature. When a conflict of multiple feature types is detected at the same geographic coordinate point, the layer is resampled based on the feature classification priority rule, and the layer data of the different types with spatial topological relationships are output.

9. A server system, characterized in that, Includes a server, the server being used to perform the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Knowledge graph construction and intelligent question and answer method and device based on deep learning

    CN119691135A

  • Transformer substation misoperation mode mining method for large-scale historical data

    CN119885095A