Natural resource element cross-layer index association method and system based on AI assistance
Through the AI-based cross-layer indicator correlation method of natural resource elements, the multimodal land feature correlation model is used to process remote sensing images and natural resource element information, which solves the efficiency and accuracy problems of remote sensing image processing in traditional methods and realizes the effective correlation and data support of cross-layer indicators.
Patent Information
- Application Number
- CN202510695186.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-05-28
AI Technical Summary
Traditional methods are unable to efficiently and accurately process massive remote sensing images and complex natural resource element information. Existing AI methods lack accuracy in feature extraction and association and model versatility, and are unable to fully explore the potential relationships between natural resource elements.
An AI-based cross-layer indicator association method for natural resource elements is adopted. Remote sensing images are processed through a multimodal land feature association model to extract the descriptive features of remote sensing image features and natural resource element information, calculate the matching degree, perform layer separation and feature association, and construct a cross-layer indicator association model.
It achieves effective linkage of natural resource elements across layers, provides an accurate data foundation, offers a scientific basis for natural resource management and decision-making, and improves the accuracy and versatility of the model.
Smart Images

Figure CN120689745A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular to an AI-assisted method and system for associating cross-layer indicators of natural resource elements. Background Art
[0002] In the field of natural resource management, accurately correlating factor indicators across different layers is crucial for rational resource planning and conservation. Traditional methods struggle to efficiently and accurately process massive amounts of remote sensing imagery and complex natural resource element information. With the advancement of AI technology, the use of AI to assist in correlating natural resource element indicators across multiple layers is becoming a trend. However, some existing AI-based methods lack accuracy in feature extraction and correlation, as well as model versatility, and are unable to fully explore the potential relationships between natural resource elements. Summary of the Invention
[0003] The purpose of the present invention is to provide an AI-assisted method and system for associating cross-layer indicators of natural resource elements.
[0004] In a first aspect, an embodiment of the present invention provides an AI-assisted method for associating natural resource elements across layers of indicators, including:
[0005] Acquire a remote sensing image set, where the remote sensing image set includes a plurality of remote sensing images;
[0006] Loading the remote sensing image set into a pre-trained multimodal land feature association model, wherein the multimodal land feature association model is used to process the plurality of remote sensing images to obtain remote sensing image features of each remote sensing image;
[0007] Loading the natural resource element information to be matched into the multimodal land feature element association model to obtain descriptive features corresponding to the natural resource element information, the descriptive features including benchmark descriptive features, land feature features, and element indicator features;
[0008] Calculating a matching degree between a descriptive feature corresponding to the natural resource element information and remote sensing image features of the plurality of remote sensing images, and determining a target remote sensing image that matches the natural resource element information based on the matching degree;
[0009] Separating the target remote sensing image into layers to obtain different types of layer data;
[0010] Extracting corresponding layer features from each layer data, and establishing a correlation relationship between each layer feature and the element indicator features in the natural resource element information;
[0011] Based on the established association relationships, a cross-layer indicator association model is constructed.
[0012] In a second aspect, an embodiment of the present invention provides a server system, including a server, wherein the server is configured to execute the method described in the first aspect.
[0013] Compared to existing technologies, the present invention provides the following beneficial effects: Using the AI-assisted cross-layer indicator association method and system disclosed in the present invention, a collection of multiple remote sensing images is acquired, and these images and the natural resource element information to be matched are loaded into a pre-trained multimodal feature association model. Remote sensing image features and descriptive features are obtained, and the target remote sensing image is determined by calculating the matching degree. The target image layer is separated to obtain different layer data, and layer features are extracted and associated with element indicator features. A cross-layer indicator association model is constructed based on this, achieving effective association of cross-layer indicators of natural resource elements. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly describes the drawings required for use in the embodiments. It should be understood that the following drawings illustrate only certain embodiments of the present invention and should not be construed as limiting the scope of the present invention. Those skilled in the art can, without inventive effort, derive other relevant drawings from these drawings.
[0015] Figure 1 A schematic diagram of the steps of the AI-assisted natural resource element cross-layer indicator association method provided in an embodiment of the present invention;
[0016] Figure 2 A schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0017] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more apparent, the technical solutions of the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings of the embodiments of the present invention. It should be understood that the described embodiments are only a portion of the embodiments of the present invention, not all of them. Generally, the components of the embodiments of the present invention described and illustrated in the drawings herein may be arranged and designed in a variety of different configurations.
[0018] The specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0019] In order to solve the technical problems in the above background technology, Figure 1 A flow chart of the AI-assisted natural resource element cross-layer indicator correlation method provided in the embodiment of the present disclosure is provided below. The AI-assisted natural resource element cross-layer indicator correlation method is introduced in detail.
[0020] Step S201: obtaining a remote sensing image set, wherein the remote sensing image set includes a plurality of remote sensing images;
[0021] Step S202: loading the remote sensing image set into a pre-trained multimodal land feature association model, wherein the multimodal land feature association model is used to process the plurality of remote sensing images to obtain remote sensing image features of each remote sensing image;
[0022] Step S203: Loading the natural resource element information to be matched into the multimodal land feature element association model to obtain descriptive features corresponding to the natural resource element information, wherein the descriptive features include a baseline descriptive feature, land feature features, and element index features;
[0023] Step S204, calculating the matching degree between the description features corresponding to the natural resource element information and the remote sensing image features of the plurality of remote sensing images, and determining a target remote sensing image that matches the natural resource element information based on the matching degree;
[0024] Step S205, performing layer separation on the target remote sensing image to obtain different types of layer data;
[0025] Step S206: extracting corresponding layer features from each layer data, and establishing a correlation between each layer feature and the element indicator features in the natural resource element information;
[0026] Step S207: construct a cross-layer indicator association model based on the established association relationship.
[0027] In an embodiment of the present invention, illustratively, the server establishes data connections with multiple satellites or multiple ground remote sensing devices. For example, these satellites may be high-resolution series satellites specifically used for natural resource monitoring. They operate at different orbital altitudes and can obtain remote sensing images of different resolutions and different bands. Ground remote sensing equipment may be distributed in specific areas, such as mountainous areas, near rivers, etc., to supplement the deficiencies of satellite remote sensing images and obtain more detailed local information. The server receives remote sensing image data from these satellites and ground devices at predetermined time intervals or according to specific trigger conditions. These data are transmitted and stored in a specific format, such as the common GeoTIFF format, which contains not only the pixel information of the image, but also important metadata such as geographic coordinates and projection information. Over time, the server gradually accumulates a large number of remote sensing images, which together constitute a remote sensing image collection. For example, during a one-month monitoring period, the server collects remote sensing images on a 24-hour basis and obtains 10 4Thousands of remote sensing images from different time periods and regions cover the various terrain, landforms, and natural resource distribution within the monitoring area. Some of these images may clearly show the distribution of large forests, while others highlight the direction of rivers and the area of water, providing a rich data foundation for subsequent analysis. After the server receives the remote sensing image collection, it loads it into a pre-trained multimodal feature association model. For example, a specific natural resource survey project focuses on land use types, vegetation cover, and mineral resource distribution in a specific area. This pre-trained multimodal feature association model enables in-depth analysis of the input remote sensing images. The model's training data comes from a large amount of historical remote sensing images and their corresponding detailed natural resource feature information. This data has been carefully annotated and organized to cover a wide range of feature types and scenarios. When the server loads the remote sensing image collection into the model, the model begins its work. It first performs feature extraction on each remote sensing image. For example, given a remote sensing image depicting mountainous terrain, the model will identify morphological features such as mountain outlines and valley orientations. It will also analyze the image's spectral information to obtain the reflectance characteristics of different features in various bands. Through a series of complex algorithms and neural network structures, the model integrates and abstracts these features, ultimately generating unique remote sensing image features for each remote sensing image. These features, represented as vectors, contain rich information about the features in the image, such as their type, distribution, and texture, providing the basis for subsequent matching with natural resource element information. In practical natural resource management, specific natural resource element information often needs to be matched and analyzed with existing remote sensing imagery. For example, a natural resource management department needs to understand the distribution of a specific type of mineral resource within a specific area. This generates natural resource element information to be matched. The server loads this matching natural resource element information into the multimodal feature association model. For example, in the case of the aforementioned mineral resource, the matching natural resource element information may include the type of mineral, its expected approximate distribution range, and a description of relevant geological features. After receiving this information, the model begins processing it. It extracts benchmark description features, land feature features and element indicator features from this information through complex semantic analysis and feature extraction algorithms.Taking mineral resources as an example, baseline descriptive features might be a general description of the mineral resource, such as "ferrous metal minerals," providing a basic semantic framework for subsequent matching. Landform features might include geological and geomorphological characteristics commonly associated with the mineral resource, such as specific rock types and mountain orientations. These features facilitate matching based on the morphology of the remote sensing image. Factor indicator features might consist of the mineral resource's classification code combined with specific indicator association rules, such as a specific mineral code and its associated reserve and grade indicators. These features provide a quantitative basis for matching. By extracting these features, the natural resource element information to be matched is converted into descriptive features comparable to the remote sensing image features, laying the foundation for the subsequent matching calculations. After obtaining the descriptive features corresponding to the natural resource element information and the remote sensing image features of each remote sensing image in the collection, the server begins calculating the degree of match between them. For example, suppose the natural resource element information is a detailed description of a forest area, including tree species, approximate area, and the topography and landforms of the forest. The remote sensing image collection includes multiple images of the region and its surrounding areas. The server uses a specific algorithm to calculate the matching degree. For example, for the semantic components of the baseline descriptive features and remote sensing image features, a word vector similarity calculation method can be used to compare the semantic information contained in the remote sensing image features in the semantic space of the description "forest" to calculate the semantic matching degree. For ground feature features, a morphological algorithm can be used to calculate the morphological matching degree by comparing the expected forest morphology (such as shape and boundaries) with the actual morphology of the ground features in the remote sensing image. For feature indicator features, such as forest area, indicators such as area obtained through image processing and analysis of the remote sensing image are compared to calculate the indicator matching degree. The matching degrees of these different dimensions are comprehensively weighted to obtain an overall matching degree. The server then selects the remote sensing images based on this matching degree. For example, by comparing the matching degree of all remote sensing images with the natural resource feature information, the remote sensing image with the highest matching degree is selected as the target remote sensing image. In this forest example, if a remote sensing image has a high degree of match with the forest feature information to be matched in terms of semantics, morphology, and indicators, then this image will be identified as the target remote sensing image. It is likely to accurately reflect the actual situation of the forest, providing an accurate data foundation for subsequent in-depth analysis. After determining the target remote sensing image, the server next separates its layers to obtain different types of layer data. Suppose the target remote sensing image is a comprehensive image covering various landforms such as cities, farmland, and rivers. The server first performs pixel-level semantic segmentation on the target remote sensing image based on a pre-trained land feature classification model. This pre-trained land feature classification model can label each pixel in the image with the corresponding land feature category.For example, pixels in urban areas are labeled as "urban buildings," pixels in farmland areas are labeled as "arable land," and pixels in river areas are labeled as "water bodies." This generates an initial layer set containing multiple categories of ground features. The server then performs spectral feature analysis and spatial topology verification on this initial layer set. Regarding spectral feature analysis, taking a specific agricultural monitoring scenario as an example, for a crop layer, the server calculates the spectral reflectance ratio of each pixel in the red-edge band to the near-infrared band. Because crops with different health conditions have different reflectance ratios in these two bands, this method can be used to monitor crop growth. Furthermore, a dynamic threshold segmentation algorithm is used to remove regions with anomalous spectral responses, such as those caused by sensor noise or local interference. For spatial topology verification, morphological closing operations are used to address fragmented patches. For example, in an urban building layer, there may be small fragmented patches due to image resolution or other reasons. Morphological closing operations can be used to merge these small patches with the surrounding main patches, making the urban building area more complete. Voronoi polygons are then used to construct a spatial adjacency matrix between features. This matrix records the spatial relationships between different features, such as which farmland is adjacent to rivers, or which urban areas are connected to roads. Finally, the server performs spatial overlay analysis on the validated layer data based on feature type. For example, when a conflict between multiple feature types is detected at the same geographic coordinate point, such as an area labeled as both farmland and urban buildings, the layer is resampled based on feature classification priority rules. If, in this project, urban buildings have a higher priority than farmland, the area is reclassified according to the urban building category. The resulting output is layer data of different types with spatial topological relationships. This layer data provides a clear and accurate data foundation for subsequent layer feature extraction and association establishment. After obtaining the different types of layer data, the server performs feature extraction on each layer and establishes associations between each layer feature and the feature indicator features in the natural resource element information. For example, in a land use monitoring project, suppose multiple layer data layers have been separated, such as cultivated land, forest land, and water bodies. For each layer, the server performs multi-scale feature extraction. For example, for a cultivated land layer, at a small scale, one might focus on detailed features such as the boundaries and texture of a single field; at a large scale, one might focus on macro features such as the distribution pattern of the entire cultivated land area and its relative position to surrounding landforms. This multi-scale feature extraction allows for comprehensive acquisition of the layer features corresponding to each layer data. These features are represented in the form of vectors or matrices and contain rich spatial and semantic information. Next, the server calculates the cross-modal similarity between the layer features and the feature indicator features.For example, in the aforementioned land use monitoring project, factor indicator features may include indicators such as cultivated land area and yield. For the cultivated land area indicator, the server quantitatively compares the regional range characteristics of the cultivated land layer features with the area indicator and calculates the similarity between them. For the yield indicator, similarity may be calculated using a specific algorithm based on the relationship between layer features such as soil type and irrigation conditions and the yield indicator. Based on these cross-modal similarities, the server performs cluster analysis. It groups layer features and factor indicator features with high similarity, thereby determining the correlation between each layer feature and the factor indicator features in the natural resource element information. For example, if the cultivated land layer features in a particular area show high similarity with factor indicator features associated with high yields in terms of soil fertility and irrigation water sources, then a correlation between cultivated land and high yields can be established, providing strong support for subsequent analysis and decision-making. After determining the correlation between each layer feature and the factor indicator features, the server begins to build a cross-layer indicator correlation model. Taking a complex integrated natural resource management scenario as an example, assume that relationships have been established between multiple layers, such as land, water resources, and vegetation, and corresponding factor indicator features (such as land quality, water flow, and vegetation coverage). First, the server performs multimodal feature fusion on these relationships. It integrates the relationships between features from different layers and the factor indicator features to generate a correlation feature matrix containing a joint spatial-semantic representation. This matrix not only contains information about the spatial distribution of each layer feature, but also includes semantic information about their relationships with the factor indicator features. For example, the specific correlation values between land layer features and land quality indicators in a particular region, as well as their corresponding spatial locations. Then, a multi-head attention mechanism is used to model cross-layer interactions within the correlation feature matrix. This mechanism captures the dynamic correlation weights between features from different layers. For example, when analyzing the relationship between water resources and vegetation, the model can dynamically adjust the correlation weights between the water layer and the vegetation layer based on the actual conditions in different regions, thereby more accurately reflecting their interactions. The server then adaptively aggregates the dynamic correlation weights with the factor indicator features in a hierarchical manner. Based on the relationships between different feature indicators and layers, the relevant information is hierarchically integrated to generate an interpretable set of indicator association rules. For example, for an ecological and environmental assessment of a specific region, a rule might be generated: when water resource flow reaches a certain threshold and vegetation coverage is within a certain range, the land quality indicator will be in good condition. The indicator association rule set is then optimized through a comparative learning strategy. Positive samples consist of cross-layer feature pairs associated with the same geographic entity, such as the layer feature pairs of land and vegetation in the same area, which are closely related in reality. Negative samples consist of layer feature pairs with conflicting spatial distributions, such as the layer feature pairs of water bodies and dry land that should not coexist in a given area.Through this comparative learning process, the indicator association rule set is continuously adjusted and optimized to become more accurate and reliable. Finally, a cross-layer indicator association model, including a multi-layer perception network, is constructed based on the optimized indicator association rule set. This model can comprehensively analyze and predict the various input layer data and feature indicator characteristics, providing a scientific basis for natural resource management and decision-making. For example, in future resource planning, by inputting different land use planning scenarios as layer data and combining them with existing feature indicator characteristics, the model can predict the impact of these scenarios on natural resources, helping decision-makers develop more reasonable plans.
[0028] In an embodiment of the present invention, the multimodal land feature association model is obtained in the following manner and can be implemented through the following examples.
[0029] The sample data set is loaded into the FLAVA model based on collaborative training to obtain the remote sensing image features, benchmark description features, land feature features and element index features of multiple training instance combinations included in the sample data set. Each training instance combination includes a remote sensing image and natural resource element information of the remote sensing image. The multimodal land feature element association model is used to: perform land feature element analysis on the natural resource element information to obtain a land feature element set of the training instance combination, generate land feature attribute information and element index information of the land feature elements in the land feature element set of the training instance combination, and generate the land feature attribute information and element index information of the land feature elements in the land feature element set of the training instance combination based on the training instance combination. The method comprises the following steps: obtaining the benchmark description features of the training instance combination based on the natural resource element information of the training instance combination, obtaining the feature features of the training instance combination based on the feature attribute information of the feature elements in the feature element set of the training instance combination, obtaining the feature index features of the training instance combination based on the feature index information of the feature elements in the feature element set of the training instance combination, and obtaining the remote sensing image features based on the remote sensing image of the training instance combination, wherein the feature attribute information is used to characterize the morphological features of the feature elements, and the feature index information is composed of the feature classification code of the feature elements combined with the indicator association rule;
[0030] Calculate the feature matching degree of any two training instance combinations based on the remote sensing image features, benchmark description features, ground object features and element index features of the plurality of training instance combinations;
[0031] Obtaining the element similarity between any two training instance combinations from the plurality of training instance combinations;
[0032] The element consistency error is calculated based on the feature matching degree and element similarity of the combination of any two training instances, and the multimodal ground feature element association model is updated based on the element consistency error.
[0033] In an embodiment of the present invention, for example, a server obtains a sample dataset containing a large amount of data from different time periods and regions within a natural resource monitoring area. Each training instance combination consists of a remote sensing image and its corresponding natural resource element information. The server loads the sample dataset into a FLAVA model based on collaborative training. For example, in one training instance combination, the remote sensing image is of a mountainous area, and the natural resource element information describes the forestland, mineral resources, and other resources within that mountainous area. The model performs feature element parsing on the natural resource element information, identifying features such as forestland and minerals, and forming a feature element set. For forestland, feature attribute information is generated, such as morphological characteristics such as shape and area, as well as feature indicator information, such as forestland type codes combined with indicator association rules such as forest cover. Baseline descriptive features, such as "distribution of natural resources in a certain mountainous area," are obtained from the overall natural resource element information. Feature features are derived based on the feature attribute information, and feature indicator features are derived based on the feature indicator information. Simultaneously, remote sensing image features are extracted from the remote sensing image. Through this process, the server obtains various features for multiple training instance combinations. The server uses the features of these training instance combinations to calculate the feature matching between any two training instance combinations. For example, a training instance combination consisting of a mountainous area is compared with another training instance combination consisting of plain farmland. The server calculates the semantic matching between the remote sensing image features and the baseline descriptor features to determine the semantic similarity between the descriptions of "mountainous area" and "plain farmland." The server also calculates the morphological matching between the remote sensing image features and the ground feature features to compare morphological differences, such as the shape of mountainous terrain and plain farmland. The server also calculates the index matching between the remote sensing image features and the feature index features, such as the matching between the forest cover index in mountainous areas and the crop yield index in plain farmland. Through these calculations, the server comprehensively measures the feature matching between the two training instance combinations. The server uses specific methods to obtain the feature similarity between any two training instance combinations. For example, the server calculates the results of a spatial overlay analysis of the number of feature elements in the feature feature sets of the two training instance combinations to examine the overlap in feature distribution. It also performs feature aggregation analysis to understand the spatial clustering characteristics of feature elements. Using these two analysis results, the server calculates the spatial coupling coefficient, resulting in the feature element coupling degree, which serves as the feature similarity measure. For example, for mountainous areas and plain farmland, the spatial distribution and aggregation relationship between forestland and farmland is analyzed to determine feature similarity. The server calculates feature consistency error based on feature matching and feature similarity. For example, the error between feature similarity and semantic, morphological, and index matching vectors is calculated separately. These errors are then combined to form a target error, which is then used to calculate the feature consistency error. A large feature consistency error indicates that there is a significant deviation between the model's current feature matching of the training instance combination and the actual feature similarity.Based on this element consistency error, the server updates the multimodal land feature association model and adjusts the model parameters so that it can more accurately process and match the relevant features of natural resource elements, improve model performance, and better associate cross-layer indicators of natural resource elements in subsequent practical applications.
[0034] In the embodiment of the present invention, the obtaining of the element similarity between any two training instance combinations among the multiple training instance combinations can be implemented through the following example.
[0035] Based on the ground feature element sets of the plurality of training instance combinations, the ground feature element coupling degree of any two training instance combinations is calculated, and the ground feature element coupling degree is used as the element similarity of the any two training instance combinations.
[0036] In this embodiment of the present invention, for example, assume that the training example combinations processed by the server are derived from monitoring data for natural resources in different regions of a province. One training example combination, A, corresponds to remote sensing imagery and related natural resource element information for the northern mountainous region of the province, with a feature set including features such as coniferous forests and granite veins. Another training example combination, B, corresponds to the central plains of the province, with a feature set including features such as farmland and irrigation canals. Based on the feature sets of these training example combinations, the server calculates the feature coupling between any two training example combinations, using this as the feature similarity. Specifically, the server first performs a spatial overlay analysis. For training example combinations A and B, the server overlays the feature distribution layer for the mountainous region with the feature distribution layer for the plains. During this process, the server determines the spatial overlap between the coniferous forests in the mountainous region and the farmland and irrigation canals in the plains, and calculates information such as the area and shape of the overlapping regions. For example, if a small area at the border between the mountainous region and the plains is both the edge of the coniferous forest in the mountainous region and close to an irrigation canal on the plains, this overlapping area will be recorded. Next, the server performs feature aggregation analysis. For coniferous forests in mountainous areas, the server analyzes their spatial aggregation characteristics, such as whether they are concentrated and continuous or scattered in small patches. Similarly, for farmland in the plains, the server analyzes their aggregation characteristics, determining whether they are large, regular fields or relatively scattered, small patches. By comparing the aggregation characteristics of the two features, the server understands the similarities and differences in their spatial organization. Then, based on the results of the spatial overlay analysis and feature aggregation analysis, the server calculates the spatial coupling coefficient. For example, if the spatial overlap between coniferous forests in mountainous areas and farmland in plains is small and their aggregation characteristics differ significantly, the calculated spatial coupling coefficient is low, indicating low feature coupling. Conversely, if there is significant overlap at the boundary and the aggregation characteristics are somewhat similar, the spatial coupling coefficient is high, indicating high feature coupling. Finally, the server uses the calculated feature coupling coefficient as the feature similarity between training instance combinations A and B. By performing such calculations on all training instance combinations, the server can comprehensively obtain the feature similarity between any two training instance combinations in multiple training instance combinations, providing key data support for the subsequent calculation of feature consistency errors and updating of multimodal land feature association models.
[0037] In the embodiment of the present invention, the calculation of the coupling degree of the ground feature elements of any two training instance combinations based on the ground feature element set of the plurality of training instance combinations can be implemented through the following examples.
[0038] Calculating a spatial overlay analysis result of the number of features in the feature set of the first training instance combination and the feature set of the second training instance combination;
[0039] Calculating an element aggregation analysis result of the number of ground feature elements in the ground feature element set of the first training instance combination and the ground feature element set of the second training instance combination;
[0040] Calculate the spatial coupling coefficient of the spatial overlay analysis result and the element aggregation analysis result of the ground feature element set of the first training instance combination and the ground feature element set of the second training instance combination to obtain the ground feature element coupling degree of the first training instance combination and the second training instance combination.
[0041] In an embodiment of the present invention, for example, a server processes natural resource monitoring data from a specific region to construct a model. Based on the feature sets of multiple training instance combinations, the server calculates the feature coupling degree between any two training instance combinations. Assume that the first training instance combination is from the coastal region of the region, and its feature set includes features such as beaches, reefs, and offshore aquaculture areas; the second training instance combination is from the inland river valley region, and its feature set includes features such as farmland, rivers, and villages. The server first performs a spatial overlay analysis on the feature sets of these two training instance combinations. The server accurately overlays the feature distribution data for the coastal region with the feature distribution data for the inland river valley region according to geographic coordinates. During this process, the server calculates the spatial overlap of different feature features. For example, if a beach in the coastal region overlaps with farmland in the inland river valley region in a small area on the map, the server records the number of features for both the beach and farmland within the overlapping area. For the beach, the server might record the percentage of the beach area within the overlapping area, converting this into the corresponding feature number. For the farmland, the server similarly records the number of features corresponding to the area within the overlapping area. Through comprehensive and detailed statistical analysis, the server obtains spatial overlay analysis results for the number of features in the first and second training instance combinations, clearly demonstrating the spatial overlap between the two features. Next, the server performs a feature aggregation analysis on the feature counts of the two training instance combinations. For coastal beaches, the server analyzes their spatial clustering characteristics, such as whether the beaches are concentrated along a long stretch of coastline or dispersed across multiple smaller beaches. For offshore aquaculture areas, the server analyzes whether large-scale, concentrated aquaculture is practiced, or whether small-scale, dispersed aquaculture is practiced. Similarly, for cultivated land in inland river valleys, the server analyzes whether it is large, continuous tracts or fragmented by rivers and roads. For villages, the server analyzes whether they are concentrated, large villages, or dispersed clusters of smaller villages. The server quantifies these clustering characteristics, calculating metrics such as concentration and dispersion index, to obtain cluster analysis data for each feature. Combining these data yields feature aggregation analysis results for the number of features in the first and second training instance combinations, reflecting the spatial organization characteristics of the features in both locations. Finally, the server calculates the spatial coupling coefficient based on the results of the spatial overlay analysis and feature aggregation analysis. The calculation of the spatial coupling coefficient comprehensively considers the degree of spatial overlap of the features and the differences in their aggregation characteristics. For example, if there is little spatial overlap between the beaches in the coastal area and the cultivated land in the inland river valley, and the concentrated distribution pattern of the beaches differs significantly from the dispersed or concentrated distribution of the cultivated land, then the calculated spatial coupling coefficient will be low, indicating that the features of these two training examples are poorly coupled.Conversely, if there's significant overlap in certain areas and the features' clustering characteristics are somewhat similar, the spatial coupling coefficient is high, and so is the degree of feature coupling. Through this calculation, the server determines the feature coupling degree for the first and second training instance combinations, providing crucial data support for the subsequent construction of a multimodal feature association model.
[0042] In the embodiment of the present invention, the calculation of the element consistency error based on the feature matching degree and element similarity of the combination of any two training instances can be implemented through the following examples.
[0043] Generate a matching degree vector of the full combination size based on the feature matching degree of the combination of any two training instances;
[0044] Generate a feature similarity vector of the full combination size based on the feature similarity of the combination of any two training instances;
[0045] The element consistency error is calculated based on the element similarity vector and the matching degree vector.
[0046] In an exemplary embodiment of the present invention, assume that a server is processing natural resource data from a large ecological reserve to construct a multimodal feature association model. This process requires calculating feature consistency errors. The server has obtained multiple training instance combinations. For example, training instance combination A corresponds to the wetland portion of the reserve, including wetland vegetation, water features, and related descriptions; training instance combination B corresponds to the nearby forest area, including various tree species, streams, and other features and descriptions. The server calculates the feature matching between any two training instance combinations, such as the semantic matching between the remote sensing image features of A and B and the baseline description features, the morphological matching between the remote sensing image features and the feature features, and the index matching between the remote sensing image features and the feature index features. After calculating the feature matching between all two training instance combinations, the server generates a matching vector for all combinations based on these results. Assume that there are five training instance combinations from different regions within the reserve, with a total of ten possible pairwise matching scenarios. The server then organizes the semantic, morphological, and index matching scores for each combination into vector form. For example, for the combination (A, B), the semantic match is 0.6, the morphological match is 0.5, and the index match is 0.4. These values are recorded sequentially at the corresponding positions in the match vector. In this way, the match information for all 10 combinations is recorded in the match vector, which comprehensively reflects the degree of feature matching between different training instance combinations. The server also calculates the feature similarity of any two training instance combinations, for example, by calculating the feature coupling between the feature sets of A and B to obtain their feature similarity. For the five training instance combinations within the protected area, there are also 10 possible pairwise combinations. The server organizes the feature similarity of each combination into a vector. For example, the feature similarity of the combination (A, B) is calculated to be 0.35, and the server records this value at the corresponding position in the feature similarity vector. This method completes the feature similarity recording for all combinations, forming a feature similarity vector for the entire combination, which reflects the degree of feature similarity between different training instance combinations. The server calculates the feature consistency error based on the generated feature similarity vector and the match vector. The server will compare the semantic, morphological, and indicator matching vectors in the matching vector with the feature similarity vector respectively. For example, first calculate the difference between the feature similarity vector and the semantic matching vector. By calculating the distribution difference of the row dimension element association distribution of the two, that is, comparing the distribution of each combination in feature similarity and semantic matching; then calculate the sum of the distribution differences of the elements of the row set of all combinations to obtain the first distribution difference. In the same way, calculate the distribution difference of the column dimension element association distribution and the sum of the distribution differences of the elements of the row set of all combinations to obtain the second distribution difference. Take the average of these two distribution differences to obtain the error between the feature similarity vector and the semantic matching vector. In the same way, calculate the error between the feature similarity vector and the morphological matching vector and the indicator matching vector.Finally, these three errors are summed to obtain the target error, based on which the feature consistency error is calculated. This error reflects the degree of consistency between feature matching and feature similarity, providing an important basis for the server to subsequently update the multimodal feature association model, thereby optimizing the model to better associate natural resource features.
[0047] In an embodiment of the present invention, the matching degree vector includes: a semantic matching degree vector, a morphological matching degree vector and an index matching degree vector. The feature matching degree of any two training instance combinations is calculated based on the remote sensing image features, benchmark description features, land feature features and element index features of the multiple training instance combinations, which can be implemented through the following examples.
[0048] Calculating the semantic matching degree between the remote sensing image features and the benchmark description features of any two training instance combinations among the plurality of training instance combinations;
[0049] Calculating the morphological matching degree between the remote sensing image features and the ground object features of any two training instance combinations among the plurality of training instance combinations;
[0050] Calculating the index matching degree between the remote sensing image features and the element index features of any two training instance combinations among the multiple training instance combinations;
[0051] Generating a matching degree vector of the full combination size based on the feature matching degree of the combination of any two training instances includes:
[0052] Generating a semantic matching degree vector of the full combination size based on the semantic matching degree of the combination of any two training instances;
[0053] Generating a morphological matching degree vector of a full combination size based on the morphological matching degrees of the combination of any two training instances;
[0054] An indicator matching degree vector of the full combination size is generated based on the indicator matching degrees of the combination of any two training instances.
[0055] In an embodiment of the present invention, for example, a server obtains multiple training instance combinations. For example, training instance combination X corresponds to a city center, containing features such as skyscrapers and commercial plazas, as well as related natural resource information. Its baseline descriptor might be "urban core business district." Training instance combination Y corresponds to a suburban farmland area, containing features such as cultivated land and irrigation facilities. Its baseline descriptor might be "suburban agricultural planting area." The server calculates the semantic match between the remote sensing image features and the baseline descriptor features of any two of the multiple training instance combinations. For combination (X, Y), the server compares the semantic information contained in the remote sensing image of the city center, such as building functions and distribution, with the baseline descriptor "urban core business district." Simultaneously, the server compares the semantic information of the remote sensing image of the suburban farmland area with the baseline descriptor "suburban agricultural planting area." Using natural language processing technology and image semantic analysis algorithms, the server quantifies the semantic similarity between the two, yielding a semantic match. For example, the calculated semantic match for combination (X, Y) is 0.2, because the semantic differences between the city center business district and the suburban agricultural planting area are significant. The server performs this calculation for all pairs of training examples. For the same training example pair X and Y, the server calculates the morphological match between their remote sensing image features and ground feature features. In remote sensing images of urban centers, high-rise buildings appear tall and dense, while commercial plazas have a specific shape and layout. In contrast, cultivated land in suburban farmland appears more regular and blocky, while irrigation facilities have a unique linear form. Using image processing and pattern recognition techniques, the server compares the degree of conformity between ground feature features in remote sensing images of urban centers and features such as high-rise buildings and commercial plazas, and the degree of conformity between ground feature features in remote sensing images of suburban farmland and features such as cultivated land and irrigation facilities. For example, by calculating the similarity of morphological parameters such as shape, size, and spatial distribution, the morphological match for the pair (X, Y) is 0.3, because the morphology of urban buildings and farmland is significantly different. This calculation is repeated for all pairs of training examples. Next, the server calculates the index match between remote sensing image features and feature index features for any two pairs of training example combinations. For example, in a training example combination X, the key indicator features of a commercial plaza might include commercial area and pedestrian flow, while the key indicator features of suburban farmland might include cultivated land area and crop yield. The server extracts information related to these indicators from remote sensing imagery. For example, it estimates the commercial plaza's area through image analysis and estimates pedestrian flow based on vehicle and pedestrian image features, comparing these information with the actual key indicators of the commercial plaza. For suburban farmland, it estimates cultivated land area and vegetation cover from imagery and correlates these with crop yield indicators. The server then calculates the degree of match between the two, yielding the indicator matching degree. For example, the indicator matching degree for the combination (X, Y) is calculated to be 0.1, due to the significant difference between the commercial and agricultural indicators.After completing the above calculations, the server generates a semantic matching vector for the entire combination based on the semantic matching between any two training instance combinations. Assuming there are eight different training instance combinations in the region, there are 28 possible pairwise combinations. The server records the semantic matching for each combination in turn into a vector, forming a semantic matching vector. Similarly, a morphological matching vector is generated based on the morphological matching, and an indicator matching vector is generated based on the indicator matching. These vectors comprehensively and systematically record the degree of matching between different training instance combinations in terms of semantics, morphology, and indicators, providing a critical data foundation for the subsequent calculation of feature consistency errors and the optimization of multimodal feature association models.
[0056] In the embodiment of the present invention, the calculation of the element consistency error based on the element similarity vector and the matching degree vector can be implemented through the following examples.
[0057] Calculating the semantic consistency error between the element similarity vector and the semantic matching vector;
[0058] Calculating the morphological consistency error between the element similarity vector and the morphological matching vector;
[0059] Calculating the indicator consistency error between the element similarity vector and the indicator matching vector;
[0060] The sum of the semantic consistency error, the morphological consistency error and the index consistency error is calculated to obtain a target error, and the element consistency error is calculated based on the target error.
[0061] In an embodiment of the present invention, illustratively, the server has generated a feature similarity vector and a semantic matching vector. For example, in the watershed data, the training instance combination M represents the mountain forest land upstream of the river, and the training instance combination N represents the plain farmland downstream of the river. The feature similarity vector records the feature similarity of all pairwise combinations such as (M, N), and the semantic matching vector records the semantic matching of the corresponding combination. For the combination (M, N), the feature similarity is calculated by the coupling degree of the ground feature elements, etc., and is assumed to be 0.25. In terms of semantic matching, the "forest ecosystem" benchmark description of the mountain forest land and the "agricultural planting area" benchmark description of the plain farmland have a semantic matching degree of 0.15 after preliminary calculation. The server calculates the semantic consistency error between the feature similarity vector and the semantic matching vector. First, calculate the distribution difference of the row dimension feature correlation distribution of the two, that is, compare the distribution of each combination in feature similarity and semantic matching. For the combination (M, N), analyze its position in the feature similarity vector and the semantic matching vector and the difference in corresponding values. The sum of the distribution differences of all row set elements in the combination is then calculated to obtain the first distribution difference. Similarly, the distribution difference of the column-dimensional feature correlation distribution and the sum of the distribution differences of all row set elements are calculated to obtain the second distribution difference. The average of these two distribution differences is taken to obtain the semantic consistency error for the combination (M, N). This calculation is repeated for all combinations, resulting in the semantic consistency error between the feature similarity vector and the semantic matching vector. Taking the same training example combinations M and N as an example, in terms of morphological matching, the remote sensing image of the mountain forest, which shows dense trees and undulating terrain, matches the forest feature, while the remote sensing image of the plain farmland, which shows large, regular farmland, matches the farmland feature. The previously calculated morphological matching degree is 0.2. The server calculates the morphological consistency error between the feature similarity vector and the morphological matching vector. Similar to the steps for calculating the semantic consistency error, the differences in the correlation distribution of the row and column dimensions for each combination in the feature similarity vector and the morphological matching vector are compared. The corresponding sum of the distribution differences is calculated and averaged to obtain the morphological consistency error for the combination (M, N). Traverse all combinations and obtain the morphological consistency error between the feature similarity vector and the morphological matching vector. For the training instance combination M, its feature indicator characteristics may include forest coverage, timber reserves, etc.; the feature indicator characteristics of the training instance combination N include cultivated land area, grain output, etc. Assume that the indicator matching degree of the combination (M, N) calculated in the previous step is 0.1. The server calculates the indicator consistency error between the feature similarity vector and the indicator matching vector. Still following the method of calculating semantic and morphological consistency errors, the distribution differences of each combination in the feature similarity vector and the indicator matching vector are compared from the row and column dimensions. The sum of the distribution differences is calculated and averaged to obtain the indicator consistency error of the combination (M, N), and then the indicator consistency error between the feature similarity vector and the indicator matching vector is obtained.The server adds the semantic consistency error, morphological consistency error, and index consistency error to obtain the target error. For example, if the semantic consistency error is 0.3, the morphological consistency error is 0.25, and the index consistency error is 0.35, the target error is 0.9. Based on this target error, the feature consistency error is calculated using a specific algorithm. This feature consistency error reflects the consistency between feature similarity and the matching degree of different dimensions in the model. The server can use this information to adjust and optimize the multimodal feature association model, improving the accuracy of the model's natural resource feature association analysis.
[0062] In an embodiment of the present invention, the error between the element similarity vector and the multimodal matching vector is calculated in the following manner:
[0063] Calculating the distribution difference between the element similarity vector and the row dimension element association distribution of the multimodal matching vector, wherein the multimodal matching vector is the semantic matching vector, the morphological matching vector, or the indicator matching vector;
[0064] Calculating the sum of distribution differences of the element similarity vector and the elements of the full combination row set of the multimodal matching vector to obtain a first distribution difference;
[0065] Calculating the distribution difference between the element similarity vector and the column dimension element association distribution of the multimodal matching vector;
[0066] Calculating the sum of the distribution differences of the element similarity vector and the elements of the full combination row set of the multimodal matching vector to obtain a second distribution difference;
[0067] An average value of the first distribution difference and the second distribution difference is calculated to obtain an error between the element similarity vector and the multimodal matching vector.
[0068] In an embodiment of the present invention, for example, assume that a server is processing data for a large region encompassing a variety of landforms and natural resource types. This region includes different areas, such as forests, grasslands, lakes, and cities, each corresponding to a training instance combination. Taking a forest region (training instance combination A) and a grassland region (training instance combination B) as examples, assume that the error between the feature similarity vector and the semantic match vector is first calculated. The feature similarity vector records the feature similarity between all training instance combinations, while the semantic match vector records the semantic match between the corresponding pairwise combinations. In the row dimension, for the combination (A, B), the feature similarity is assumed to be 0.3, indicating the overall feature similarity between the forest and grassland; the semantic match is assumed to be 0.1, reflecting the similarity in their semantic descriptions. The server analyzes the relative position and size differences between these two values in their respective vector rows and quantifies these differences using a specific algorithm (such as calculating the absolute value or relative ratio of the difference) to obtain the distribution difference of the feature correlation distribution in the row dimension. This calculation is performed for all training instance combinations (all combinations). The server sums the distribution differences calculated for each combination in the row dimension to obtain the first distribution difference. For example, if there are 10 training instance combinations in the entire region, there are 45 possible pairwise combinations. The server sums the distribution differences for each of these 45 combinations in the row dimension. Assuming the sum is 12.5, this is the first distribution difference. This value reflects the overall distribution difference between the feature similarity vector and the semantic match vector in the row direction. Taking the combination (A, B) as an example, in the column dimension, the server again analyzes the relative position and size differences between the feature similarity 0.3 and the semantic match 0.1 in their respective vector columns. This difference is quantified using an algorithm similar to the row dimension to obtain the distribution difference of the feature association distribution in the column dimension. This calculation is repeated for all combinations in the entire region. Similar to the calculation of the first distribution difference, the server sums the distribution differences calculated for each combination in the column dimension to obtain the second distribution difference. Assume that in the column dimension, the sum of the distribution differences of these 45 combinations is 10.5. This is the second distribution difference, which reflects the overall distribution difference between the feature similarity vector and the semantic matching vector in the column direction. The server adds the first distribution difference of 12.5 and the second distribution difference of 10.5, and then divides it by 2 to get an average of 11.5. This average is the error between the feature similarity vector and the semantic matching vector. In this way, the server can accurately measure the degree of difference between the feature similarity vector and the semantic matching vector. The same steps are also applied to calculate the error between the feature similarity vector and the morphological matching vector, and the feature similarity vector and the indicator matching vector. These errors provide an important basis for the server to further optimize the multimodal feature association model, which helps to improve the accuracy of the model's analysis of the association of natural resource features across layers.
[0069] In an embodiment of the present invention, the calculation of the feature matching degree of any two training instance combinations based on the remote sensing image features, benchmark description features, ground object features and element index features of the multiple training instance combinations can be implemented through the following examples.
[0070] Calculating the integrated descriptive features of each training instance combination based on the baseline descriptive features, the feature features, and the element index features of each training instance combination;
[0071] Calculating a feature matching degree between the remote sensing image features and the integrated description features of any two training instance combinations among the plurality of training instance combinations;
[0072] The calculating of the element consistency error based on the feature matching degree and the correlation index matching degree of the combination of any two training instances includes:
[0073] Generate a matching degree vector of the full combination size based on the feature matching degree of the combination of any two training instances;
[0074] Calculating the element similarity of the arbitrary two training instance combinations based on the correlation index matching degree of the arbitrary two training instance combinations, and generating an element similarity vector of the full combination size based on the element similarity of the arbitrary two training instance combinations;
[0075] The element consistency error is calculated based on the element similarity vector and the matching degree vector.
[0076] In an embodiment of the present invention, for example, assume that a server is processing natural resource data for a national nature reserve. The reserve encompasses a variety of ecological regions, including mountains, forests, lakes, and grasslands, with each region corresponding to a training instance combination. For each training instance combination, taking the mountain region as an example, its baseline descriptive feature might be "high-altitude mountains with rich mineral resources." Feature features include topographical characteristics such as mountain shape and slope, while factor indicator features include specific indicators such as mineral type and mountain altitude. The server calculates an integrated descriptive feature based on these features. Using a specific algorithm, for example, the baseline descriptive feature is semantically encoded, the feature features and factor indicator features are numerically processed, and then these features are fused according to certain weights. Assume that the baseline descriptive feature has a weight of 0.3, the feature feature has a weight of 0.4, and the factor indicator feature has a weight of 0.3. The semantically encoded and numerically converted features are multiplied by their corresponding weights and then added together to obtain the integrated descriptive feature for the mountain region training instance combination. The integrated descriptive feature for all training instance combinations within the reserve is calculated in the same manner. Next, the server calculates the feature matching between the remote sensing image features and the integrated descriptive features of any two training instance combinations from the multiple training instance combinations. For example, consider a comparison of training instance combinations from a mountainous region and a forested region. The remote sensing image features of the mountainous region reflect information such as the mountain's topography and vegetation cover, while the integrated descriptive features of the forested region contain comprehensive information such as its semantics, features, and feature indicators. The server uses a similarity calculation algorithm to match the image's texture, spectrum, and spatial structure with the various components of the integrated descriptive features. For example, the vegetation cover in the mountainous remote sensing image is compared with the vegetation-related content in the integrated descriptive features of the forest. By calculating the degree of similarity between the two, the feature matching between the two training instance combinations is calculated. This calculation is performed for all pairs of training instance combinations. Based on the feature matching between any two training instance combinations, the server generates a matching vector for all combinations. Assume that there are eight training instance combinations from different ecological regions within the protected area, with 28 possible pairwise combinations. The server records the feature matching between each combination in turn in a vector, forming a matching vector that comprehensively records the feature matching between different combinations. The server calculates feature similarity based on the matching between the associated indicators of any two training instance combinations. For example, there's a correlation between mineral indicators in mountainous areas and timber resource indicators in forested areas. By analyzing the matching of these correlated indicators and combining methods such as feature coupling, feature similarity is calculated. Assume the feature similarity for the mountain and forest combination is 0.25. After calculating feature similarity for all combinations, a feature similarity vector covering all combinations is generated. Again, using eight training example combinations as an example, the feature similarities for all 28 combinations are recorded in a vector. Finally, the server calculates feature consistency error based on the feature similarity vector and the matching vector.Using the previously described method of calculating and averaging the row and column distribution differences, the error between the two is calculated. For example, the first distribution difference is calculated for the row dimension element association distribution, and then the sum of the distribution differences for all combined row elements is calculated to obtain the first distribution difference. Similarly, the second distribution difference is calculated for the column dimension correlation value, and the two are averaged to obtain the element consistency error. This error reflects the consistency between feature matching and element similarity in the model. The server can use this information to optimize the multimodal feature association model and improve the accuracy of cross-layer indicator correlation analysis of natural resource elements within nature reserves.
[0077] In the embodiment of the present invention, the calculation of the integrated description features of each training instance combination based on the benchmark description features, the feature of the ground object and the feature index features of each training instance combination can be implemented through the following examples.
[0078] The benchmark description features, land feature features and element index features of each training instance combination are subjected to multimodal element balanced aggregation or multimodal element weighted aggregation to obtain the integrated description features of each training instance combination.
[0079] In this embodiment of the present invention, for example, assume that the server processes data for a large region encompassing a variety of geographical environments and natural resource types. This region includes different areas such as mountains, plains, rivers, and wetlands, with each region constituting a training instance combination. Taking the mountain training instance combination as an example, its baseline descriptive feature is "high-altitude mountainous area rich in metal minerals." Feature features include topographical characteristics such as mountain morphology and rock texture. Factor indicator features include specific numerical indicators such as mineral reserves and mountain altitude. When performing balanced aggregation of multimodal features, the server assigns equal weight to these three features. First, the baseline descriptive features are subjected to text vectorization, converting textual information into numerical vectors. For example, word embedding technology is used to convert "high-altitude mountainous area rich in metal minerals" into a vector of a specific dimension. For feature features, image processing techniques are used to quantize information such as mountain morphology and rock texture into numerical vectors. The factor indicator features are numerical in nature and can be used directly. The server then performs a simple averaging operation on these three numerical vectors. For example, assuming the baseline descriptive feature vector is [0.2, 0.3, 0.4], the feature vector is [0.1, 0.4, 0.3], and the factor indicator feature vector is [0.3, 0.2, 0.5]. The integrated descriptive feature vector obtained through balanced aggregation is [(0.2 + 0.1 + 0.3) / 3, (0.3 + 0.4 + 0.2) / 3, (0.4 + 0.3 + 0.5) / 3], or [0.2, 0.3, 0.4]. The server then performs multimodal balanced aggregation on all training instance combinations in the region in the same manner, obtaining the integrated descriptive features for each combination. Using the mountainous training instance combination as an example, given that the region primarily focuses on mineral resource development, the server assigns a higher weight of 0.5 to the factor indicator feature, a weight of 0.3 to the baseline descriptive feature, and a weight of 0.2 to the feature feature. Similarly, the baseline descriptive feature text is first vectorized, the feature features are digitized, and the factor indicator features remain in numerical form. Assume that the baseline descriptive feature vector is [0.2, 0.3, 0.4], the feature vector is [0.1, 0.4, 0.3], and the feature index vector is [0.3, 0.2, 0.5]. Through weighted aggregation, the integrated descriptive feature vector is [(0.2 × 0.3 + 0.1 × 0.2 + 0.3 × 0.5), (0.3 × 0.3 + 0.4 × 0.2 + 0.2 × 0.5), (0.4 × 0.3 + 0.3 × 0.2 + 0.5 × 0.5)], or [0.23, 0.27, 0.43]. The server flexibly adjusts weights based on the characteristics of different training instance combinations and analysis requirements, performing multimodal feature weighted aggregation on all training instance combinations in the region to obtain integrated descriptive features for each combination. These integrated descriptive features integrate multiple aspects of information, laying the foundation for subsequent calculation of feature matching between training instance combinations and optimization of the multimodal feature association model.
[0080] In an embodiment of the present invention, before updating the multimodal ground feature element association model based on the element consistency error, the following implementation is also provided.
[0081] Calculating the standard element matching degree of any two training instance combinations based on the standard element labels of the multiple training instance combinations;
[0082] Calculating a comparison error based on the feature matching degree and the standard element matching degree of the combination of any two training instances;
[0083] The updating of the multimodal ground feature element association model based on the element consistency error includes:
[0084] Calculating a final error of the multimodal ground feature element association model based on the element consistency error and the comparison error;
[0085] The multimodal land feature association model is updated based on the final error.
[0086] In an embodiment of the present invention, for example, a server is processing natural resource data for a large ecological park encompassing multiple ecological regions, such as forests, wetlands, and farmland. Each region corresponds to a training instance combination, which is used to construct a multimodal feature association model. The server obtains standard feature labels for multiple training instance combinations. For example, the standard feature labels for the forest region training instance combination specify information such as tree species types and coverage standard ranges for the forest in that region; the standard feature labels for the wetland region training instance combination specify information such as water quality standards and biodiversity indicators for the wetland. Based on these standard feature labels, the server calculates the standard feature matching degree between any two training instance combinations. Taking the forest and wetland training instance combinations as examples, the server compares the tree species types of the forest with the tree species suitable for wetland growth (if relevant standards exist), as well as the forest coverage standard range and the impact standards of wetland ecology on surrounding vegetation cover. Using a specific algorithm, the server quantifies the degree of matching between the two training instance combinations in terms of standard feature labels, resulting in a standard feature matching degree. For example, the calculated standard feature matching degree for the forest and wetland training instance combinations is 0.2, indicating a low degree of similarity in standard features between the two training instance combinations. The server performs this calculation for all combinations of two training instances. The server has calculated the feature matching degree for any combination of two training instances. For example, for the forest and wetland training instance combination, the feature matching degree based on feature features such as remote sensing image features and benchmark descriptor features has been previously calculated to be 0.3. The server calculates the comparative error based on the feature matching degree and the standard feature matching degree. By analyzing the difference between the feature matching degree and the standard feature matching degree, a specialized error calculation method is used, such as calculating the sum of the squares of the differences and then taking the average. Assume that the comparative error for the forest and wetland training instance combination is 0.05. This calculation is performed for all combinations of two training instances to obtain the overall comparative error. The server has also calculated the feature consistency error. For example, by combining the previously calculated errors between the feature similarity vector and the semantic, morphological, and index matching vectors, the server determines a feature consistency error of 0.1. The server calculates the final error of the multimodal feature association model based on the feature consistency error and the comparative error. Assuming a simple addition method (more complex calculation methods may be used depending on the model's characteristics), we add the element consistency error of 0.1 and the average contrast error (assuming it is 0.08) obtained by aggregating the contrast errors, resulting in a final error of 0.18. The server updates the multimodal feature association model based on this final error. The model contains a series of parameters, such as weights in a neural network. The server uses optimization algorithms, such as stochastic gradient descent, to adjust these parameters based on the final error.This enables the model to reduce the difference between feature matching and standard element matching, as well as the inconsistency between element similarity and each matching vector, when processing a similar training instance combination next time. This improves the accuracy and reliability of the model's correlation analysis of cross-layer indicators of natural resource elements in ecological parks, and better supports the management and planning of ecological parks.
[0087] In the embodiment of the present invention, the calculation of the contrast error based on the feature matching degree and the standard element matching degree of the combination of any two training instances can be implemented through the following examples.
[0088] Generate a matching degree vector of the full combination size based on the feature matching degree of the combination of any two training instances;
[0089] Generate a standard element identification vector of the full combination size based on the standard element matching degree of the combination of any two training instances, wherein the standard element matching degrees on the self-correlation diagonal of the elements in the standard element identification vector are set as valid flags, and the remaining standard element matching degrees are reset to invalid flags;
[0090] The comparison error is calculated based on the matching degree vector and the standard element identification vector.
[0091] In an embodiment of the present invention, for example, assume that a server is processing data from a large nature reserve encompassing multiple ecological regions, such as grasslands, forests, and lakes, with each ecological region corresponding to a training instance combination. The server has calculated the feature matching degree between any two training instance combinations. For example, for the grassland and forest training instance combination, by comparing their remote sensing image features, benchmark descriptor features, ground feature features, and element index features, the feature matching degree is 0.35; the feature matching degree between the grassland and lake training instance combination is 0.2, and so on. The server generates a matching degree vector for all combinations based on the feature matching degrees of any two training instance combinations. Assume that there are five training instance combinations from different ecological regions within the reserve, with a total of 10 possible pairwise combinations. The server sequentially records the feature matching degrees of these 10 combinations into a vector, forming a matching degree vector. For example, the matching degree vector might be [0.35, 0.2, 0.4, 0.15, 0.3, 0.25, 0.1, 0.45, 0.28, 0.32], which comprehensively records the degree of feature matching between different training instance combinations. Each training instance combination has a corresponding standard feature label. For example, the standard feature labels for the grassland training instance combination include grass species type and vegetation cover standard; the standard feature labels for the forest training instance combination include tree species type and canopy density standard. The server generates a standard feature identification vector for the entire combination based on the standard feature matching scores between any two training instance combinations. During this generation process, standard feature matching scores on the self-correlation diagonal (i.e., the matching score between a training instance combination and itself) are set to valid. For example, the standard feature matching score between grassland and grassland itself is set to 1 (indicating a valid match), and the standard feature matching score between forest and forest itself is also set to 1. Other non-self-correlation standard feature matching scores, such as those between grassland and forest, grassland and lake, etc., are reset to invalid and set to 0. Assume that the standard feature identification vector generated for five training instance combinations is [1, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 1]; this vector highlights the self-correlation standard feature matching. The server calculates the comparison error based on the matching degree vector and the standard feature identification vector. The server uses a specific algorithm, such as calculating the sum of the squares of the differences between the corresponding elements of the two vectors. The server subtracts the corresponding elements of the matching degree vector [0.35, 0.2, 0.4, 0.15, 0.3, 0.25, 0.1, 0.45, 0.28, 0.32] from the corresponding elements of the standard feature identification vector [1, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 1, 0, 0, 0, 0, 0, 1], squares them, and sums them. This calculation yields a numerical value, which is the comparison error.This comparison error reflects the degree of difference between the feature matching degree and the standard element matching degree, providing an important basis for the server to further update the multimodal land feature association model to improve the accuracy of the model's cross-layer indicator association analysis of natural resource elements in nature reserves.
[0092] In the embodiment of the present invention, the layer separation of the target remote sensing image and the acquisition of different types of layer data can be implemented through the following examples.
[0093] Performing pixel-level semantic segmentation on the target remote sensing image based on a pre-trained ground feature classification model to generate an initial layer set containing multiple categories of ground features;
[0094] Performing spectral feature analysis and spatial topology verification on the initial layer set, wherein the spectral feature analysis includes calculating the spectral reflectance ratio of each pixel in the red-edge band and the near-infrared band, and eliminating abnormal spectral response areas based on a dynamic threshold segmentation algorithm; the spatial topology verification includes using morphological closing operations to process fragmented patches and constructing a spatial adjacency matrix between elements using Voronoi polygons;
[0095] The verified layer data is subjected to spatial overlay analysis according to the type of land feature. When conflicts of multiple feature types are detected at the same geographic coordinate point, the layer is resampled based on the feature classification priority rules to output the different types of layer data with spatial topological relationships.
[0096] In an embodiment of the present invention, for example, a server is processing a target remote sensing image of a large area encompassing multiple landform types, such as cities, rural areas, farmland, forests, and rivers. The server aims to separate the target remote sensing image into layers to obtain different types of layer data for subsequent analysis of natural resource elements. The server uses a pre-trained landform classification model to perform pixel-level semantic segmentation on the target remote sensing image. This landform classification model, trained on a large amount of accurately labeled remote sensing image data, can accurately identify various landform elements. For example, in this target remote sensing image, the model classifies pixels in urban areas as "buildings," pixels in farmland areas as "cultivated land," pixels in forest areas as "vegetation," and pixels in river areas as "water bodies." By classifying each pixel in the image, an initial layer set containing multiple landform elements is generated. Each layer in this set corresponds to a landform category, such as "building layer," "cultivated land layer," "vegetation layer," "water body layer," and so on. Each layer records the pixel distribution of the corresponding landform element in the image. The server performs spectral signature analysis on the initial layer set. Taking the "vegetation layer" as an example, spectral signature analysis involves calculating the ratio of the spectral reflectance of each pixel in the red-edge band to the near-infrared band. Because vegetation of different health and types has different reflectance ratios in these two bands, for example, healthy green vegetation has higher reflectance in the near-infrared band and relatively lower reflectance in the red-edge band, calculating this ratio effectively distinguishes different vegetation conditions. Furthermore, the server uses a dynamic threshold segmentation algorithm to remove areas with abnormal spectral response. During image acquisition, some pixels may exhibit abnormal spectral response due to sensor noise, atmospheric interference, or local conditions. The dynamic threshold segmentation algorithm dynamically determines a threshold based on the overall spectral characteristics of the image. Pixel areas with spectral reflectance ratios outside this threshold are identified as having abnormal spectral response and are removed. For example, in the "vegetation layer," the spectral reflectance ratios of some pixels significantly deviate from the normal vegetation range, possibly due to local shadows or sensor failure. These areas are identified and removed by the algorithm, thereby improving the quality of the layer data. The server also performs spatial topology verification on the initial layer set. On the one hand, morphological closing operations are used to handle fragmented patches. In the "Building Layer," due to image resolution limitations or other factors, there may be some isolated, fragmented small patches. These small patches may be local details of buildings that have been separated due to resolution issues and are not independent features. The morphological closing operation merges these fragmented small patches with the surrounding adjacent main patches through an operation of dilation followed by erosion, making the building boundaries more continuous and complete, consistent with the actual spatial form. On the other hand, the server constructs a spatial adjacency matrix between features using Voronoi polygons.Taking the "cultivated land layer" and "road layer" as examples, the Voronoi polygon algorithm divides the space into multiple polygonal regions based on the location of each feature. Points within each polygon have the shortest distance to the corresponding feature. This method clearly identifies the spatial adjacency relationships between different features. For example, which cultivated land is adjacent to roads, which forests border rivers, and so on. These relationships are recorded in a matrix format. This spatial adjacency matrix provides an important basis for subsequent analysis of the interactions and spatial layout of features. The server performs spatial overlay analysis on the validated layer data based on feature type. During this analysis, conflicts between multiple feature types may be detected at the same geographic coordinate point. For example, in a certain area, the "building layer" and the "cultivated land layer" may overlap at some coordinate points, meaning that the area is identified as both buildings and cultivated land. In this case, the server resamples the layer based on the feature classification priority rules. Assume that the feature classification priority rules for this project give buildings a higher priority than cultivated land. Then, for conflicting coordinate points, the server will redefine the area according to the building category, classify it into the "building layer", and discard the identification of cultivated land. Through this process, the conflict of feature types is eliminated. Ultimately, the server outputs different types of layer data with spatial topological relationships. These layer data not only accurately reflect the distribution of different land features, but also contain topological information such as the spatial adjacency relationship between them, laying a solid data foundation for the subsequent extraction of layer features, establishing associations with natural resource element information, and building cross-layer indicator association models.
[0097] In an embodiment of the present invention, the extraction of corresponding layer features for each layer data and the establishment of an association relationship between each layer feature and the element indicator features in the natural resource element information can be implemented through the following examples.
[0098] Performing multi-scale feature extraction on each layer data to obtain layer features corresponding to each layer data;
[0099] Calculating the cross-modal similarity between the layer features and the element indicator features;
[0100] A cluster analysis is performed based on the cross-modal similarity to determine the correlation between the features of each layer and the feature indicator features in the natural resource feature information.
[0101] In an embodiment of the present invention, for example, assume that a server is processing remote sensing image data for a large agricultural region. This region contains multiple features, such as cultivated land, irrigation facilities, and roads, and the corresponding natural resource element information is known, including feature indicators such as crop yield and irrigation water consumption. The server performs multi-scale feature extraction on each layer of data. Taking the cultivated land layer as an example, at a small scale, the server focuses on the detailed features of individual fields. For example, using high-resolution image information, the server extracts the clarity of field boundaries and the texture characteristics of the soil within the fields. These small-scale features can reflect the meticulous management of the farmland. For example, the regularity of field boundaries may indicate the rationality of farmland planning. At a large scale, the server focuses on the macroscopic characteristics of the entire cultivated land area. For example, the server analyzes the overall distribution pattern of cultivated land, determining whether it is concentrated and contiguous or dispersed, and the relative positional relationship between cultivated land and surrounding features (such as irrigation facilities and roads). Large-scale features help to understand the spatial layout of cultivated land within the entire region and its interconnections with other elements. By comprehensively extracting features at different scales, the server obtains layer features corresponding to the cultivated land layer. These features, represented as vectors or matrices, comprehensively cover all aspects of cultivated land, from microscopic to macroscopic levels. Similarly, the server performs similar multi-scale feature extraction on layers such as irrigation facilities and roads, obtaining their corresponding layer features. The server then calculates cross-modal similarity between these layer features and the feature index. For the cultivated land layer features and the feature index of crop yield, the server employs a specific algorithm. For example, given that layer features such as cultivated land area and soil quality are closely related to crop yield, the server quantifies the features reflecting this information in the cultivated land layer and compares them with the crop yield index. Suppose, through analysis, a functional relationship is found between the soil fertility feature value in a particular region's cultivated land layer and the region's crop yield. The server uses this relationship to calculate similarity between the two. For the irrigation facility layer features and the feature index of irrigation water consumption, the server analyzes the correlation between layer features such as irrigation facility coverage and water delivery capacity and irrigation water consumption. By comparing the numerical and logical connections between the two, the server derives cross-modal similarity. Similarly, cross-modal similarities are calculated for all layer features and corresponding factor indicator features. Based on the calculated cross-modal similarities, the server performs cluster analysis. For example, the server organizes the cross-modal similarity data of all cultivated land layer features and crop yields. If it is found that the cultivated land layer features of certain areas have a high similarity with the factor indicator features of high-yield crop yields in terms of soil fertility, irrigation conditions, etc., the server will classify these areas with similar features into one category. Similarly, for the irrigation facility layer features and the irrigation water consumption factor indicator features, those areas with high similarity with different water consumption demands in terms of facility coverage, water delivery efficiency, etc. are clustered separately.Through this cluster analysis, the server can clearly identify the correlations between the features of each layer and the indicator characteristics of the natural resource element information. For example, it can clarify which specific layer characteristics of cultivated land correspond to high yields, and which characteristics of irrigation facilities are associated with high irrigation water consumption. This provides a key basis for the subsequent construction of a cross-layer indicator correlation model.
[0102] In the embodiment of the present invention, the construction of a cross-layer indicator association model based on the established association relationship can be implemented through the following examples.
[0103] Performing multimodal feature fusion on the correlation between the features of each layer and the features of the element indicators to generate a correlation feature matrix containing a spatial-semantic joint representation;
[0104] Based on the multi-head attention mechanism, the cross-layer interaction modeling of the correlation feature matrix is performed to capture the dynamic correlation weights between features of different layers;
[0105] Adaptively performing hierarchical aggregation on the dynamic association weights and factor indicator features to generate an interpretable indicator association rule set;
[0106] The indicator association rule set is optimized by a contrastive learning strategy, wherein positive samples are composed of cross-layer feature pairs associated with the same geographic entity, and negative samples are composed of layer feature pairs with conflicting spatial distributions;
[0107] Based on the optimized indicator association rule set, a cross-layer indicator association model containing a multi-layer perception network is constructed.
[0108] In an embodiment of the present invention, for example, assume that the server is processing data from a large ecological conservation area. This conservation area contains various landforms, such as forests, rivers, and wetlands, corresponding to multiple layers of data. Furthermore, various natural resource factor indicator features, such as forest cover and river flow, are known. The server first performs multimodal feature fusion on the correlation between each layer feature and the factor indicator features. For example, forest layer features include information such as tree species and distribution density, while river layer features include river width and flow velocity. Forest cover and river flow are corresponding factor indicator features. The server integrates these different types of features. For the forest layer, tree species are represented as semantic vectors, and distribution density is presented as numerical values. For the river layer, width and flow velocity are also converted to numerical values. At the same time, factor indicator features such as forest cover and river flow are also incorporated. Using a specific algorithm, these features from different modalities (semantic, numerical, etc.) are combined to generate a correlation feature matrix containing a joint spatial-semantic representation. Each row of the matrix may represent a specific area, while the columns correspond to different features. This matrix not only reflects the spatial distribution of each layer feature but also incorporates semantic information and its correlation with the factor indicator features. The server uses a multi-head attention mechanism to model cross-layer interactions in the correlation feature matrix. Taking the forest and river layers as an example, the multi-head attention mechanism enables the model to focus on the relationships between the features of the forest and river layers from different perspectives. In a given area, the presence of forests may affect river water quality and flow, while rivers, in turn, provide water for the forests. Using the multi-head attention mechanism, the model analyzes these interactions across multiple dimensions, such as forest vegetation type and area, and river flow and velocity. This process captures the dynamic correlation weights between features in different layers. For example, in arid regions, the influence of forest vegetation on river flow may be more significant, while in humid regions, the weight of rivers on forest water replenishment may be more significant. This approach allows the model to dynamically determine the weights of important correlations between layers in different regions. The server adaptively aggregates these dynamic correlation weights with factor indicator features in a hierarchical manner. For example, based on the dynamic correlation weights between forests and rivers in different regions, as well as factor indicator features such as forest cover and river flow, an interpretable set of indicator association rules is generated. For areas with high forest cover and close proximity to rivers, a rule might be "Higher forest cover combined with appropriate river flow contributes to maintaining good ecological diversity." These rules are derived through hierarchical aggregation based on the interrelationships between features from different layers and feature indicators, using dynamic association weights. They clearly explain the connections and impacts between different natural resource elements. The server optimizes the indicator association rule set using a comparative learning strategy. Positive samples consist of pairs of cross-layer features associated with the same geographic entity, such as a forest and its surrounding, affected river. These pairs are ecologically related and constitute positive samples.Negative samples consist of pairs of layer features with conflicting spatial distributions, such as forests and large areas of exposed rock that are unlikely to coexist in reality (assuming such a situation does not exist in the protected area). The server uses these positive and negative samples to optimize the indicator association rule set. For example, if a rule in the rule set indicates no correlation between forest cover and river flow, but positive samples show a clear correlation, the server will adjust the rule. By continuously comparing the differences between positive and negative samples and the rule set, the rule set becomes more accurate and complete. Based on the optimized indicator association rule set, the server constructs a cross-layer indicator association model using a multi-layer perceptron network. The multi-layer perceptron network can learn and apply complex indicator association rules. For example, when new remote sensing imagery data for a region of the protected area is input, after layer separation and feature extraction, the model, based on the optimized indicator association rule set and the multi-layer perceptron network, can predict the characteristics of natural resource indicators in that area. For example, it can predict the impact of changes in forest cover on river ecology or infer the ecological status of surrounding wetlands based on changes in river flow. This model provides a powerful analytical tool for natural resource management and ecological research in the protected area.
[0109] The embodiment of the present invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned AI-assisted natural resource element cross-layer indicator association method. Figure 2 As shown, Figure 2 This is a structural block diagram of a computer device 100 provided in an embodiment of the present invention. The computer device 100 includes a memory 111 , a processor 112 , and a communication unit 113 .
[0110] To achieve data transmission or interaction, the memory 111, processor 112, and communication unit 113 are electrically connected to each other directly or indirectly. For example, these components can be electrically connected to each other via one or more communication buses or signal lines.
[0111] For illustrative purposes, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the present disclosure to the precise forms disclosed. Numerous modifications and variations are possible in light of the above teachings. These embodiments have been selected and described in order to best illustrate the principles of the present disclosure and its practical application, thereby enabling those skilled in the art to best utilize the present disclosure and to utilize various embodiments with various modifications as appropriate for the specific application contemplated.
Claims
1. The AI-assisted cross-layer indicator correlation method for natural resource elements is characterized by: include: Acquire a remote sensing image set, where the remote sensing image set includes a plurality of remote sensing images; Loading the remote sensing image set into a pre-trained multimodal land feature association model, wherein the multimodal land feature association model is used to process the plurality of remote sensing images to obtain remote sensing image features of each remote sensing image; Loading the natural resource element information to be matched into the multimodal land feature element association model to obtain descriptive features corresponding to the natural resource element information, the descriptive features including benchmark descriptive features, land feature features, and element indicator features; Calculating a matching degree between a descriptive feature corresponding to the natural resource element information and remote sensing image features of the plurality of remote sensing images, and determining a target remote sensing image that matches the natural resource element information based on the matching degree; Separating the target remote sensing image into layers to obtain different types of layer data; Extracting corresponding layer features from each layer data, and establishing a correlation relationship between each layer feature and the element indicator features in the natural resource element information; Based on the established association relationships, a cross-layer indicator association model is constructed.
2. The method according to claim 1, characterized in that The multimodal ground feature association model is obtained by: The sample data set is loaded into the FLAVA model based on collaborative training to obtain the remote sensing image features, benchmark description features, land feature features and element index features of multiple training instance combinations included in the sample data set. Each training instance combination includes a remote sensing image and natural resource element information of the remote sensing image. The multimodal land feature element association model is used to: perform land feature element analysis on the natural resource element information to obtain a land feature element set of the training instance combination, generate land feature attribute information and element index information of the land feature elements in the land feature element set of the training instance combination, and generate the land feature attribute information and element index information of the land feature elements in the land feature element set of the training instance combination based on the training instance combination. The method comprises the following steps: obtaining the benchmark description features of the training instance combination based on the natural resource element information of the training instance combination, obtaining the feature features of the training instance combination based on the feature attribute information of the feature elements in the feature element set of the training instance combination, obtaining the feature index features of the training instance combination based on the feature index information of the feature elements in the feature element set of the training instance combination, and obtaining the remote sensing image features based on the remote sensing image of the training instance combination, wherein the feature attribute information is used to characterize the morphological features of the feature elements, and the feature index information is composed of the feature classification code of the feature elements combined with the indicator association rule; Calculate the feature matching degree of any two training instance combinations based on the remote sensing image features, benchmark description features, ground object features and element index features of the plurality of training instance combinations; Calculating a spatial overlay analysis result of the number of features in the feature set of the first training instance combination and the feature set of the second training instance combination; Calculating an element aggregation analysis result of the number of ground feature elements in the ground feature element set of the first training instance combination and the ground feature element set of the second training instance combination; Calculating the spatial coupling coefficient of the spatial overlay analysis results and the feature aggregation analysis results of the feature feature set of the first training instance combination and the feature feature set of the second training instance combination to obtain the feature feature coupling degree of the first training instance combination and the second training instance combination, and using the feature feature coupling degree as the feature similarity between any two training instance combinations; Generate a matching degree vector of the full combination size based on the feature matching degree of the combination of any two training instances; Generate a feature similarity vector of the full combination size based on the feature similarity of the combination of any two training instances; An element consistency error is calculated based on the element similarity vector and the matching degree vector, and the multimodal ground feature element association model is updated based on the element consistency error.
3. The method according to claim 2, characterized in that The matching degree vector includes: a semantic matching degree vector, a morphological matching degree vector and an index matching degree vector. The feature matching degree of any two training instance combinations is calculated based on the remote sensing image features, benchmark description features, ground feature features and element index features of the plurality of training instance combinations, including: Calculating the semantic matching degree between the remote sensing image features and the benchmark description features of any two training instance combinations among the plurality of training instance combinations; Calculating the morphological matching degree between the remote sensing image features and the ground object features of any two training instance combinations among the plurality of training instance combinations; Calculating the index matching degree between the remote sensing image features and the element index features of any two training instance combinations among the multiple training instance combinations; Generating a matching degree vector of the full combination size based on the feature matching degree of the combination of any two training instances includes: Generating a semantic matching degree vector of the full combination size based on the semantic matching degree of the combination of any two training instances; Generating a morphological matching degree vector of a full combination size based on the morphological matching degrees of the combination of any two training instances; An indicator matching degree vector of the full combination size is generated based on the indicator matching degrees of the combination of any two training instances.
4. The method according to claim 3, characterized in that The calculating the element consistency error based on the element similarity vector and the matching degree vector includes: Calculating the semantic consistency error between the element similarity vector and the semantic matching vector; Calculating the morphological consistency error between the element similarity vector and the morphological matching vector; Calculating the indicator consistency error between the element similarity vector and the indicator matching vector; The sum of the semantic consistency error, the morphological consistency error and the index consistency error is calculated to obtain a target error, and the element consistency error is calculated based on the target error.
5. The method according to claim 4, characterized in that The error between the element similarity vector and the multimodal matching vector is calculated as follows: Calculating the distribution difference between the element similarity vector and the row dimension element association distribution of the multimodal matching vector, wherein the multimodal matching vector is the semantic matching vector, the morphological matching vector, or the indicator matching vector; Calculating the sum of the distribution differences of the element similarity vector and the elements of the full combination row set of the multimodal matching vector to obtain a first distribution difference; Calculating the distribution difference between the element similarity vector and the column dimension element association distribution of the multimodal matching vector; Calculating the sum of the distribution differences of the element similarity vector and the elements of the full combination row set of the multimodal matching vector to obtain a second distribution difference; An average value of the first distribution difference and the second distribution difference is calculated to obtain an error between the element similarity vector and the multimodal matching vector.
6. The method according to claim 2, characterized in that The calculating of the feature matching degree of any two training instance combinations based on the remote sensing image features, benchmark description features, ground object features, and element index features of the plurality of training instance combinations includes: Performing multimodal element balanced aggregation or multimodal element weighted aggregation on the benchmark descriptive features, feature features, and element index features of each training instance combination to obtain an integrated descriptive feature of each training instance combination; Calculating a feature matching degree between the remote sensing image features and the integrated description features of any two training instance combinations among the plurality of training instance combinations; The calculating of the element consistency error based on the feature matching degree and the correlation index matching degree of the combination of any two training instances includes: Generate a matching degree vector of the full combination size based on the feature matching degree of the combination of any two training instances; Calculating the element similarity of the arbitrary two training instance combinations based on the correlation index matching degree of the arbitrary two training instance combinations, and generating an element similarity vector of the full combination size based on the element similarity of the arbitrary two training instance combinations; The element consistency error is calculated based on the element similarity vector and the matching degree vector.
7. The method according to claim 2, characterized in that Before updating the multimodal ground feature element association model based on the element consistency error, the method further includes: Calculating the standard element matching degree of any two training instance combinations based on the standard element labels of the multiple training instance combinations; Generate a matching degree vector of the full combination size based on the feature matching degree of the combination of any two training instances; Generate a standard element identification vector of the full combination size based on the standard element matching degree of the combination of any two training instances, wherein the standard element matching degrees on the self-correlation diagonal of the elements in the standard element identification vector are set as valid flags, and the remaining standard element matching degrees are reset to invalid flags; Calculating a comparison error based on the matching degree vector and the standard element identification vector; The updating of the multimodal ground feature element association model based on the element consistency error includes: Calculating a final error of the multimodal ground feature element association model based on the element consistency error and the comparison error; The multimodal land feature association model is updated based on the final error.
8. The method according to claim 1, characterized in that The step of separating the target remote sensing image into layers and obtaining different types of layer data includes: Performing pixel-level semantic segmentation on the target remote sensing image based on a pre-trained ground feature classification model to generate an initial layer set containing multiple categories of ground features; Performing spectral feature analysis and spatial topology verification on the initial layer set, wherein the spectral feature analysis includes calculating the spectral reflectance ratio of each pixel in the red-edge band and the near-infrared band, and eliminating abnormal spectral response areas based on a dynamic threshold segmentation algorithm; the spatial topology verification includes using morphological closing operations to process fragmented patches and constructing a spatial adjacency matrix between elements using Voronoi polygons; The verified layer data is subjected to spatial overlay analysis according to the type of land feature. When conflicts of multiple feature types are detected at the same geographic coordinate point, the layer is resampled based on the feature classification priority rules to output the different types of layer data with spatial topological relationships.
9. The method according to claim 1, characterized in that The extracting corresponding layer features from each layer data and establishing a correlation between each layer feature and the element indicator features in the natural resource element information includes: Performing multi-scale feature extraction on each layer data to obtain layer features corresponding to each layer data; Calculating the cross-modal similarity between the layer features and the element indicator features; Performing cluster analysis based on the cross-modal similarity to determine the correlation between each layer feature and the feature indicator features in the natural resource element information; The cross-layer indicator association model is constructed based on the established association relationship, including: Performing multimodal feature fusion on the correlation between the features of each layer and the features of the element indicators to generate a correlation feature matrix containing a spatial-semantic joint representation; Based on the multi-head attention mechanism, the cross-layer interaction modeling of the correlation feature matrix is performed to capture the dynamic correlation weights between features of different layers; Adaptively performing hierarchical aggregation on the dynamic association weights and factor indicator features to generate an interpretable indicator association rule set; The indicator association rule set is optimized by a contrastive learning strategy, wherein positive samples are composed of cross-layer feature pairs associated with the same geographic entity, and negative samples are composed of layer feature pairs with conflicting spatial distributions; Based on the optimized indicator association rule set, a cross-layer indicator association model containing a multi-layer perception network is constructed.
10. A server system, characterized in that: The method comprises a server, wherein the server is configured to execute the method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Environmental pollution source-tracing system and method based on characteristic pollution factor source analysis
CN110085281A
Intelligent monitoring method and system for typical earth surface elements guided by base vectors
CN116822980A
Image text matching method based on cross-modal interaction
CN118364302A
Natural resource asset data acquisition method and platform
CN118470550A
Intelligent site selection method and system based on domain knowledge graph, and electronic equipment
CN118627937A