Territorial space planning multi-source heterogeneous data intelligent integration system based on semantic analysis
By introducing semantic density equalization and boundary fuzzy infectious eigenvalues, combined with deep learning and CLIP model, the semantic distribution imbalance and boundary ambiguity problems caused by classification system differences in national land space planning are solved, and high-precision data fusion and calibration are achieved, improving the robustness and maintainability of the data.
Patent Information
- Application Number
- CN202510837539.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-23
AI Technical Summary
Traditional methods have failed to effectively solve the problems of semantic distribution imbalance and boundary ambiguity caused by classification system differences in multi-source heterogeneous data in land space planning, resulting in semantic conflicts and information loss during data fusion.
Semantic density equalization and boundary fuzzy infectious eigenvalues are introduced, and dynamic reliability evaluation mechanism is constructed through deep learning models, and semantic calibration is performed by combining CLIP model and knowledge graph to achieve cross-modal semantic alignment and data calibration.
It significantly improves the accuracy and consistency of multi-source data fusion, can automatically identify and calibrate semantic conflicts and boundary blur problems, form a closed-loop optimization mechanism, and improves the robustness and maintainability of data.
Smart Images

Figure CN120336940A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multi-source heterogeneous data fusion for territorial spatial planning, and specifically to an intelligent integration system for multi-source heterogeneous data of territorial spatial planning based on semantic analysis. Background Art
[0002] Territorial spatial planning involves the integration and analysis of multi-source heterogeneous data, including remote sensing images, GIS data, current land use maps, planning texts, and social and economic statistical information, etc. Traditional data integration methods mainly rely on manual annotation, rule matching, or simple statistical similarity calculation, and it is difficult to effectively solve the problem of semantic heterogeneity between multi-source data. In recent years, some studies have attempted to use natural language processing and computer vision technologies to perform semantic parsing on unstructured data, but most methods only focus on data alignment of a single modality and lack a systematic evaluation of cross-modal semantic consistency. In addition, existing technologies usually use static thresholds or fixed rules to judge data reliability and cannot dynamically adapt to the complex problems of semantic boundary ambiguity and classification system differences in territorial spatial planning, resulting in semantic conflicts or information loss during data integration.
[0003] The existing technologies have the following deficiencies:
[0004] Traditional methods do not fully consider the information distribution differences of different data sources under a unified semantic framework. For example, remote sensing images and planning texts may adopt different classification systems, resulting in local over-density or over- sparsity of semantic units after mapping, affecting the balance of data fusion. Existing technologies lack a quantitative evaluation of the semantic density balance degree and are difficult to identify the problem of semantic distribution dispersion caused by classification system differences. In territorial spatial planning, there is often semantic ambiguity in land type boundaries (such as the urban-rural fringe). Existing methods mostly judge the data consistency of boundary regions based on geometric overlap or simple semantic similarity, ignoring the "infectious" impact of fuzzy semantics on adjacent regions. For example, the gradual change region of land use types may cause semantic label leakage of adjacent data sources, and existing technologies lack the ability to dynamically model the semantic pollution radius and cannot effectively suppress the diffusion effect of boundary ambiguity.
[0005] In view of the above deficiencies, the present invention introduces the abnormal eigenvalue of semantic density balance degree and the eigenvalue of boundary fuzzy infectivity to quantify the semantic distribution dispersion degree and boundary leakage risk of multi-source data respectively. Further, a dynamic reliability evaluation mechanism is constructed by combining a deep learning model to solve the hidden impact of unbalanced cross-modal semantic distribution and the dynamic contagion effect of boundary fuzzy semantics. Summary of the Invention
[0006] The purpose of the present invention is to provide an intelligent integration system for multi-source heterogeneous data of territorial spatial planning based on semantic analysis to solve the problems in the above background.
[0007] The object of the present invention can be achieved by the following technical solutions:
[0008] An intelligent integration system for multi-source heterogeneous data in territorial space planning based on semantic analysis, comprising:
[0009] A data acquisition and recognition module, which is used to obtain territorial space planning data, and identify and select the semantic density balance degree and the boundary fuzzy contagion as the core evaluation dimensions;
[0010] A semantic density balance degree analysis module, which generates an abnormal feature value of the semantic density balance degree by calculating the information entropy difference of each data source under a unified semantic framework, and is used to evaluate the dispersion degree of the information distribution of different classification systems;
[0011] A boundary fuzzy contagion analysis module, which calculates the semantic pollution radius of the adjacent data area by analyzing the semantic gradient change of the boundary fuzzy area, generates a boundary fuzzy contagion feature value, and is used to evaluate whether there is a semantic leakage phenomenon in the territorial space planning data;
[0012] A comprehensive reliability evaluation module, which constructs a comprehensive reliability evaluation vector with the abnormal feature value of the semantic density balance degree and the boundary fuzzy contagion feature value, inputs it into a deep learning model for analysis, and divides the territorial space planning data into reliable data and unreliable data according to the analysis result;
[0013] A semantic calibration feedback module, which calibrates the territorial space planning data based on the unreliable data.
[0014] As a further solution of the present invention: the evaluation of the dispersion degree of the information distribution of different classification systems specifically includes:
[0015] According to the identified and selected semantic density balance degree, calculate the abnormal feature value of the semantic density balance degree by calculating the information entropy difference of each data source under a unified semantic framework, and judge whether the abnormal feature value of the semantic density balance degree is greater than or equal to a preset threshold. If so, the information distribution of different classification systems is dispersed. If not, the information distribution of different classification systems is concentrated.
[0016] As a further solution of the present invention: the obtaining process of the abnormal feature value of the semantic density balance degree is:
[0017] Map the semantic units in the multi-source heterogeneous national territorial space planning data to a unified semantic representation space through the CLIP model to achieve cross-modal semantic alignment; subsequently, use the kernel density estimation method combined with the Gaussian kernel function and the dynamic bandwidth selection strategy based on reinforcement learning to perform local density modeling on the semantic distributions of each data source; construct a probability distribution based on the local density function and calculate the information entropy; at the same time, construct a semantic relationship graph, use the graph attention network to perform weighted modeling on the semantic similarity between different data sources, and on this basis, improve the calculation method of the Kullback-Leibler divergence to evaluate the deviation degree of the semantic distributions between data sources; finally, combine the information entropy and the KL divergence to construct a semantic density balance degree anomaly scoring function and output the semantic density balance degree anomaly eigenvalue.
[0018] As a further solution of the present invention: evaluating whether there is a semantic leakage phenomenon in the national territorial space planning data specifically includes:
[0019] According to the identified boundary fuzzy contagion, by analyzing the semantic gradient change in the boundary fuzzy area, calculate its semantic pollution radius for the adjacent data area, generate a boundary fuzzy contagion eigenvalue, and determine whether the boundary fuzzy contagion eigenvalue is greater than or equal to a preset threshold. If so, there is a semantic leakage phenomenon in the national territorial space planning data; if not, there is no semantic leakage phenomenon in the national territorial space planning data.
[0020] As a further solution of the present invention: the process of obtaining the boundary fuzzy contagion eigenvalue is as follows:
[0021] First, identify the boundary areas between different land types through a semantic segmentation model and define them as semantic fuzzy areas; calculate the change rate of the semantic similarity gradient between adjacent land types in the fuzzy area, and use the sliding window method to extract the local semantic gradient direction and amplitude; establish a semantic pollution propagation model based on the gradient change to simulate the influence range of semantic leakage, and solve the semantic pollution radius through a diffusion equation; finally, combine the pollution radius and the semantic similarity decline rate to construct a boundary fuzzy contagion scoring function and generate a boundary fuzzy contagion eigenvalue.
[0022] As a further solution of the present invention: constructing a comprehensive reliability evaluation vector from the semantic density balance degree anomaly eigenvalue and the boundary fuzzy contagion eigenvalue and inputting it into a deep learning model for analysis specifically includes:
[0023] Obtain the semantic density balance degree anomaly eigenvalue and the boundary fuzzy contagion eigenvalue, construct a comprehensive reliability evaluation vector from the semantic density balance degree anomaly eigenvalue and the boundary fuzzy contagion eigenvalue as the input of the deep learning model, train the deep learning model, and according to the trained deep learning model, output the confidence score of the national territorial space planning data. The deep learning model selects a convolutional neural network model.
[0024] As a further solution of the present invention: The training process of the deep learning model is as follows:
[0025] Taking the comprehensive reliability evaluation vector as an input sample, input it into a convolutional neural network model for training. The convolutional neural network model consists of several one-dimensional convolutional layers, pooling layers, and fully connected layers, and is used to extract local correlation and high-order interaction information in the feature vector; during the training process, a cross-entropy loss function is adopted, and manually labeled reliable or unreliable samples are used as supervision signals. The network parameters are continuously optimized through the backpropagation algorithm to improve the accuracy of the model's semantic consistency judgment.
[0026] As a further solution of the present invention: The division of the territorial space planning data into reliable data and unreliable data specifically includes:
[0027] Set a threshold interval according to the confidence score output by the deep learning model. If the score is higher than the first threshold, it is determined as highly reliable data; if the score is between the first threshold and the second threshold, it is determined as generally reliable data; if the score is lower than the second threshold, it is determined as unreliable data.
[0028] As a further solution of the present invention: The calibration of the territorial space planning data based on the unreliable data specifically includes:
[0029] For the data marked as unreliable, call the context information in the semantic relationship graph, combine the knowledge graph reasoning mechanism to locate the most likely semantic attribution category; use the multi-modal CLIP model to re-encode the territorial space planning data and compare and match it with the benchmark vector in the standard semantic framework; adopt a transfer learning strategy to transfer the semantic mapping rules in the reliable data to the unreliable data to complete the semantic label correction; re-enter the calibrated data into the system for verification until the reliability evaluation standard is met.
[0030] The beneficial effects of the present invention:
[0031] (1) By introducing two core evaluation dimensions, namely "semantic density balance" and "boundary fuzzy contagion", the present invention constructs a systematic and multi-level semantic consistency analysis framework, which can deeply depict the semantic differences and potential conflicts between multi-source heterogeneous data in territorial spatial planning from both macro and micro levels. Among them, the semantic density balance focuses on the concentrated or discrete characteristics of the overall semantic distribution, realizes cross-modal semantic alignment with the help of the CLIP model, and combines kernel density estimation and information entropy theory to improve the traditional calculation method of Kullback-Leibler divergence, quantifying the distribution deviation degree of different data sources in the unified semantic space, so as to effectively identify the semantic conflict risk caused by the classification system difference. The boundary fuzzy contagion, on the other hand, focuses on the local semantic stability of land use class boundaries. By identifying fuzzy areas through semantic segmentation, combining semantic similarity gradient extraction and diffusion propagation model, a scoring function for semantic pollution radius and similarity decline rate is constructed to accurately evaluate whether there is semantic leakage phenomenon and its influence range in the boundary area. The above two-dimensional fusion analysis mechanism not only realizes the comprehensive perception of the semantic quality of multi-source data, but also provides key feature inputs for subsequent comprehensive reliability evaluation, significantly improving the fusion accuracy, expression consistency and application credibility of heterogeneous data in the unified semantic framework, and having strong engineering practicability and popularization value.
[0032] (2) Based on the comprehensive evaluation results of the abnormal eigenvalue of semantic density balance and the eigenvalue of boundary fuzzy contagion by the deep learning model, the system can realize the intelligent discrimination of the reliability of territorial spatial planning data and automatically identify unreliable data with problems such as semantic conflicts, distribution offsets or boundary fuzziness. Once abnormal data is found, the system will immediately trigger the semantic calibration process, use the context association information in the semantic relationship graph, and combine the semantic reasoning mechanism driven by the knowledge graph to accurately locate the source of the semantic attribution deviation. On this basis, the multi-modal CLIP model is further introduced to re-encode the original data and perform high-dimensional space matching with the benchmark vector in the standard semantic framework to improve the accuracy of semantic mapping. At the same time, the system adopts a transfer learning strategy to extract the verified semantic mapping rules from the trusted data set and adapt and transfer them to the unreliable data to realize the dynamic correction and consistency enhancement of semantic labels. The calibrated data will be re-input into the system for verification, forming a closed-loop optimization mechanism of "evaluation - identification - correction - feedback". This mechanism not only improves the adaptive ability of the system in complex semantic environments, but also significantly enhances the robustness and maintainability of territorial spatial planning data in the multi-source fusion process, providing a solid technical support for building a high-quality and sustainable updated spatial planning database. BRIEF DESCRIPTION OF THE DRAWINGS
[0033] The present invention will be further described below with reference to the accompanying drawings.
[0034] Figure 1 This is the flowchart of the intelligent integration system for multi-source heterogeneous data of territorial spatial planning based on semantic analysis of the present invention. Specific Embodiments
[0035] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0036] Please refer to Figure 1 As shown, the present invention is an intelligent integration system for multi-source heterogeneous data of territorial spatial planning based on semantic analysis, including:
[0037] A data collection and recognition module, which is used to obtain territorial spatial planning data, and identify and select the semantic density balance degree and the boundary fuzzy contagion as the core evaluation dimensions;
[0038] A semantic density balance degree analysis module, which generates an abnormal characteristic value of the semantic density balance degree by calculating the information entropy difference of each data source under a unified semantic framework, and is used to evaluate the dispersion degree of the information distribution of different classification systems;
[0039] A boundary fuzzy contagion analysis module, which calculates the semantic pollution radius of the boundary fuzzy area to adjacent data areas by analyzing the semantic gradient change of the boundary fuzzy area, and generates a boundary fuzzy contagion characteristic value, and is used to evaluate whether there is a semantic leakage phenomenon in the territorial spatial planning data;
[0040] A comprehensive reliability evaluation module, which constructs a comprehensive reliability evaluation vector with the abnormal characteristic value of the semantic density balance degree and the boundary fuzzy contagion characteristic value, and inputs it into a deep learning model for analysis. According to the analysis results, the territorial spatial planning data is divided into reliable data and unreliable data;
[0041] A semantic calibration feedback module, which calibrates the territorial spatial planning data based on the unreliable data.
[0042] In the data collection and recognition module, the data collection and recognition module is used to obtain territorial spatial planning data, and identify and select the semantic density balance degree and the boundary fuzzy contagion as the core evaluation dimensions, specifically including:
[0043] The data acquisition and recognition module is used to obtain multi-source heterogeneous data related to territorial spatial planning, and preliminarily identify and analyze the semantic features in the data. Specifically, this module realizes the automated acquisition and structured processing of the basic data required for territorial spatial planning by docking various data sources, including remote sensing image databases, geographic information system (GIS) platforms, current land use maps, planning text files, and socio-economic statistical data sets; during the data acquisition process, metadata parsing technology is used to extract the semantic description information of each type of data, and combined with natural language processing and image semantic recognition algorithms, unstructured or semi-structured data is converted into unified and comparable semantic units.
[0044] Furthermore, the data acquisition and recognition module conducts semantic consistency analysis on the collected semantic units, identifies and selects "semantic density balance degree" and "boundary fuzzy contagion" as the core evaluation dimensions for subsequent semantic quality analysis. Among them, the semantic density balance degree is used to measure the concentration or dispersion of semantic distributions under different classification systems, and the boundary fuzzy contagion is used to detect whether there is semantic leakage in the land use type boundary area; this module outputs the initial evaluation results of the two types of characteristic values and transfers them to the next analysis module to support the intelligent judgment and calibration of the overall semantic consistency of territorial spatial planning data.
[0045] In the semantic density balance degree analysis module, the semantic density balance degree analysis module generates an abnormal characteristic value of the semantic density balance degree by calculating the information entropy difference of each data source under a unified semantic framework, which is used to evaluate the dispersion degree of information distribution of different classification systems. Specifically, it includes:
[0046] According to the identified semantic density balance degree, calculate the abnormal characteristic value of the semantic density balance degree by calculating the information entropy difference of each data source under a unified semantic framework, and judge whether the abnormal characteristic value of the semantic density balance degree is greater than or equal to the preset threshold. If so, the information distribution of different classification systems is dispersed; if not, the information distribution of different classification systems is concentrated.
[0047] The process of obtaining the abnormal characteristic value of the semantic density balance degree is as follows:
[0048] Map the semantic units in the multi-source heterogeneous territorial spatial planning data to a unified semantic representation space through the CLIP model to achieve cross-modal semantic alignment; then, use the kernel density estimation method combined with the Gaussian kernel function and the dynamic bandwidth selection strategy based on reinforcement learning to perform local density modeling on the semantic distributions of each data source.
[0049] Construct a probability distribution based on the local density function and calculate the information entropy of each data source, which is used to measure the concentration or dispersion of the internal semantic distribution of each data source. The calculation expression of the information entropy is: ; where Represents the data source The information entropy of Represents the data source Represents the data source The number of semantic units contained in Represents the th data source and the th probability of the semantic unit Represents the natural logarithm Represents the th semantic unit Represents the th data source;
[0050] At the same time, construct a semantic relationship graph , and use the graph attention network to weight and model the semantic similarity between different data sources. The attention coefficient between nodes is calculated by the following formula: ; where Represents the attention coefficient between node and node , Represents a normalization function used to map each element in a vector to the interval and make the sum of these elements equal to , Represents a learnable vector used to calculate the score of the feature combination between nodes Represents the vector transpose Represents a learnable weight matrix used to linearly transform the node features Represents the feature vector of node , Represents the feature vector of node , Represents the concatenation operation, that is, connecting two vectors into a longer vector and represent two different nodes;
[0051] And on this basis, improve the calculation method of the Kullback-Leibler divergence to evaluate the semantic distribution deviation between data sources. The specific improved calculation expression is: ; where Represents the Kullback-Leibler divergence (KL divergence) from data source to data source , Represents at the sample point , the semantic weight factor between data source and , Indicates at the sample point the probability density function value of the data source at that point; represents the probability density function value of the data source ; and represent two different data sources, represents the data source sample point;
[0052] Finally, by combining information entropy and KL divergence, a semantic density balance degree anomaly scoring function is constructed to output the semantic density balance degree anomaly eigenvalue, which is used to reflect the consistency level of semantic distribution between different classification systems. Among them, the calculation expression of the semantic density balance degree anomaly scoring function is: ; where represents the semantic density balance degree anomaly eigenvalue, and are preset weight factors, represents the average information entropy of all data sources, represents the average KL divergence between any two data sources.
[0053] It should be noted that: The CLIP is a multi-modal pre-training model developed by OpenAI. It aims to map text and images to the same embedding space through contrastive learning, so as to achieve cross-modal understanding and matching.
[0054] In the boundary fuzzy contagion analysis module, the boundary fuzzy contagion analysis module calculates the semantic pollution radius of the boundary fuzzy area on the adjacent data area by analyzing the semantic gradient change in the boundary fuzzy area, and generates a boundary fuzzy contagion eigenvalue to evaluate whether there is a semantic leakage phenomenon in the territorial space planning data. Specifically, it includes:
[0055] According to the identified boundary fuzzy contagion, by analyzing the semantic gradient change in the boundary fuzzy area, calculate the semantic pollution radius of the boundary fuzzy area on the adjacent data area, generate a boundary fuzzy contagion eigenvalue, and judge whether the boundary fuzzy contagion eigenvalue is greater than or equal to the preset threshold. If so, there is a semantic leakage phenomenon in the territorial space planning data. If not, there is no semantic leakage phenomenon in the territorial space planning data.
[0056] The process of obtaining the boundary fuzzy contagion eigenvalue is as follows:
[0057] First, use a semantic segmentation model to identify the boundary area between different land classes and define it as a semantic fuzzy area; calculate the semantic similarity gradient change rate between adjacent land classes in the fuzzy area, and use the sliding window method to extract the local semantic gradient direction and amplitude; establish a semantic pollution propagation model based on the gradient change, simulate the influence range of semantic leakage, and solve the semantic pollution radius through the diffusion equation.
[0058] Finally, by combining the pollution radius and the decline rate of semantic similarity, a boundary-fuzzy infectivity scoring function is constructed to generate boundary-fuzzy infectivity eigenvalues;
[0059] Among them, the calculation expression of the decline rate of semantic similarity is: ; among them, represents the decline rate of semantic similarity, represents the average semantic similarity inside the fuzzy area, represents the average semantic similarity outside the fuzzy area, represents the distance;
[0060] The calculation expression of the boundary-fuzzy infectivity scoring function is: ; among them, represents the boundary-fuzzy infectivity eigenvalue, and represent preset weight factors, represents the semantic pollution radius, represents the decline rate of semantic similarity.
[0061] In the comprehensive reliability evaluation module, the comprehensive reliability evaluation module constructs a comprehensive reliability evaluation vector from the semantic density balance degree abnormal eigenvalue and the boundary-fuzzy infectivity eigenvalue, and inputs it into the deep learning model for analysis. According to the analysis results, the national territorial space planning data is divided into reliable data and unreliable data, specifically including:
[0062] Obtain the semantic density balance degree abnormal eigenvalue and the boundary-fuzzy infectivity eigenvalue, construct a comprehensive reliability evaluation vector from the semantic density balance degree abnormal eigenvalue and the boundary-fuzzy infectivity eigenvalue, use it as the input of the deep learning model, train the deep learning model, and according to the trained deep learning model, output the confidence score of the national territorial space planning data. The deep learning model selects a convolutional neural network model.
[0063] The training process of the deep learning model is:
[0064] Taking the comprehensive reliability evaluation vector as an input sample, input it into a convolutional neural network model for training. The convolutional neural network model consists of several one-dimensional convolutional layers, pooling layers, and fully connected layers, and is used to extract local correlations and high-order interaction information in the feature vector. During the training process, a cross-entropy loss function is adopted, and manually labeled reliable or unreliable samples are used as supervision signals. The network parameters are continuously optimized through the backpropagation algorithm to improve the accuracy of the model's semantic consistency judgment. Finally, after the model training is completed, the trained convolutional neural network model is used to predict new land spatial planning data, and the corresponding confidence score is output. The confidence score represents the probability that the land spatial planning data is judged as reliable data under the current semantic framework.
[0065] Set a threshold interval according to the confidence score output by the deep learning model. If the score is higher than the first threshold, it is determined as highly reliable data; if the score is between the first threshold and the second threshold, it is determined as generally reliable data; if the score is lower than the second threshold, it is determined as unreliable data. For unreliable data, mark it as an object to be calibrated, and record its data source, geographical area, and semantic conflict type to form a dataset to be calibrated.
[0066] In the semantic calibration feedback module, the semantic calibration feedback module calibrates the land spatial planning data based on the unreliable data, specifically including:
[0067] For the data marked as unreliable, call the context information in the semantic relationship graph, combine the knowledge graph reasoning mechanism to locate the most likely semantic attribution category; re-encode the original data using the multi-modal CLIP model and compare and match it with the benchmark vector in the standard semantic framework; adopt a transfer learning strategy to transfer the semantic mapping rules in the reliable data to the unreliable data to complete the semantic label correction; finally, re-enter the calibrated data into the system for verification until the reliability evaluation standard is met.
[0068] Working principle of the present invention: The present invention aims to solve the problem of inconsistency in classification systems, expression granularity, and semantic boundaries of multi-source data. The system includes a data collection and recognition module, a semantic density balance analysis module, a boundary fuzzy contagion analysis module, a comprehensive reliability evaluation module, and a semantic calibration feedback module. The data collection module extracts unified semantic units by docking multi-source data such as remote sensing images, GIS platforms, land use maps, and planning texts, and identifies "semantic density balance" and "boundary fuzzy contagion" as the core evaluation dimensions. The semantic density balance analysis module realizes cross-modal semantic alignment through the CLIP model, combines kernel density estimation and information entropy calculation, improves the KL divergence method to evaluate the deviation degree of semantic distribution, and constructs a semantic density balance anomaly scoring function to output eigenvalues. The boundary fuzzy contagion analysis module identifies the land use type boundary fuzzy area through semantic segmentation, extracts the semantic gradient change rate, establishes a diffusion model to solve the semantic pollution radius, and constructs a boundary fuzzy contagion scoring function to judge whether there is a semantic leakage phenomenon. The comprehensive reliability evaluation module forms an evaluation vector with the two types of eigenvalues and inputs it into a convolutional neural network model for training, outputs a confidence score, and divides the data reliability level. The semantic calibration feedback module then calls the semantic graph and knowledge reasoning mechanism based on unreliable data, combines transfer learning to complete semantic correction, and finally realizes the intelligent integration and semantic consistency optimization of land space planning data. This technical solution effectively improves the accuracy and practicability of multi-source heterogeneous data fusion, providing high-quality data support for land space planning decision-making.
[0069] The above has described in detail an embodiment of the present invention, but the content described is only the preferred embodiment of the present invention and cannot be considered as limiting the scope of implementation of the present invention. All equivalent changes and improvements made according to the scope of the application of the present invention should still fall within the scope covered by the patent of the present invention.
Claims
1. An intelligent integration system for multi-source heterogeneous data in territorial spatial planning based on semantic analysis, characterized in that, Including: A data acquisition and recognition module, which is used to obtain land spatial planning data, and identify and select the semantic density balance degree and boundary fuzzy contagion as the core evaluation dimensions; A semantic density balance degree analysis module, which generates an abnormal eigenvalue of the semantic density balance degree by calculating the information entropy difference of each data source under a unified semantic framework, and is used to evaluate the information distribution dispersion degree of different classification systems; A boundary fuzzy contagion analysis module, which calculates the semantic pollution radius of the adjacent data area by analyzing the semantic gradient change of the boundary fuzzy area, generates a boundary fuzzy contagion eigenvalue, and is used to evaluate whether there is a semantic leakage phenomenon in the land spatial planning data; A comprehensive reliability evaluation module, which constructs a comprehensive reliability evaluation vector from the abnormal eigenvalue of the semantic density balance degree and the boundary fuzzy contagion eigenvalue, and inputs it into a deep learning model for analysis. According to the analysis results, the land spatial planning data is divided into reliable data and unreliable data; A semantic calibration feedback module, which calibrates the land spatial planning data based on the unreliable data.
2. The intelligent integration system for multi-source heterogeneous data of territorial spatial planning based on semantic analysis according to claim 1, wherein The evaluation of the information distribution dispersion degree of different classification systems specifically includes: According to the identified semantic density balance degree, by calculating the information entropy difference of each data source under a unified semantic framework, calculating the abnormal eigenvalue of the semantic density balance degree, and judging whether the abnormal eigenvalue of the semantic density balance degree is greater than or equal to a preset threshold. If so, the information distribution of different classification systems is dispersed; if not, the information distribution of different classification systems is concentrated.
3. The intelligent integration system for multi-source heterogeneous data of territorial spatial planning based on semantic analysis according to claim 2, characterized in that, The process of obtaining the abnormal eigenvalue of the semantic density balance degree is as follows: Map the semantic units in the multi-source heterogeneous land spatial planning data to a unified semantic representation space through the CLIP model to achieve cross-modal semantic alignment; then use the kernel density estimation method combined with the Gaussian kernel function and the dynamic bandwidth selection strategy based on reinforcement learning to perform local density modeling on the semantic distribution of each data source; construct a probability distribution based on the local density function and calculate the information entropy; at the same time, construct a semantic relationship graph, use the graph attention network to perform weighted modeling on the semantic similarity between different data sources, and improve the calculation method of the Kullback-Leibler divergence on this basis to evaluate the semantic distribution deviation degree between data sources; finally, combine the information entropy and KL divergence to construct a semantic density balance degree abnormal scoring function and output the abnormal eigenvalue of the semantic density balance degree.
4. The intelligent integration system for multi-source heterogeneous data of territorial spatial planning based on semantic analysis according to claim 1, wherein The evaluation of whether there is a semantic leakage phenomenon in the land spatial planning data specifically includes: According to the identified boundary fuzzy contagion, by analyzing the semantic gradient change of the boundary fuzzy area, calculating the semantic pollution radius of the adjacent data area, generating a boundary fuzzy contagion eigenvalue, and judging whether the boundary fuzzy contagion eigenvalue is greater than or equal to a preset threshold. If so, there is a semantic leakage phenomenon in the land spatial planning data; if not, there is no semantic leakage phenomenon in the land spatial planning data.
5. The intelligent integration system for multi-source heterogeneous data of territorial space planning based on semantic analysis according to claim 4, characterized in that, The process of obtaining the boundary fuzzy contagion eigenvalue is as follows: First, identify the boundary regions between different land classes through a semantic segmentation model and define them as semantic ambiguity zones; calculate the rate of change of semantic similarity gradient between adjacent land classes within the ambiguity zones, and use the sliding window method to extract the local semantic gradient direction and amplitude; establish a semantic pollution propagation model based on the gradient change, simulate the influence range of semantic leakage, and solve the semantic pollution radius through a diffusion equation; finally, combine the pollution radius and the rate of decrease of semantic similarity to construct a boundary ambiguity contagion scoring function and generate boundary ambiguity contagion eigenvalues.
6. The intelligent integration system for multi-source heterogeneous data of territorial spatial planning based on semantic analysis according to claim 1, characterized in that, Construct a comprehensive reliability evaluation vector with the abnormal eigenvalue of semantic density balance degree and the boundary ambiguity contagion eigenvalue, and input it into a deep learning model for analysis, specifically including: Obtain the abnormal eigenvalue of semantic density balance degree and the boundary ambiguity contagion eigenvalue, construct a comprehensive reliability evaluation vector with the abnormal eigenvalue of semantic density balance degree and the boundary ambiguity contagion eigenvalue as the input of the deep learning model, train the deep learning model, and according to the trained deep learning model, output the confidence score of the national territorial space planning data. The deep learning model selects a convolutional neural network model.
7. The intelligent integration system for multi-source heterogeneous data of territorial spatial planning based on semantic analysis according to claim 6, wherein The training process of the deep learning model is as follows: Take the comprehensive reliability evaluation vector as the input sample and input it into the convolutional neural network model for training. The convolutional neural network model consists of several one-dimensional convolutional layers, pooling layers and fully connected layers, and is used to extract the local correlation and high-order interaction information in the feature vector; during the training process, use the cross-entropy loss function, use the manually labeled reliable or unreliable samples as the supervision signal, and continuously optimize the network parameters through the backpropagation algorithm to improve the accuracy of the model's judgment of semantic consistency.
8. The intelligent integration system for multi-source heterogeneous data of territorial spatial planning based on semantic analysis according to claim 1, characterized in that The division of the national territorial space planning data into reliable data and unreliable data specifically includes: Set a threshold interval according to the confidence score output by the deep learning model. If the score is higher than the first threshold, it is determined as highly reliable data; if the score is between the first threshold and the second threshold, it is determined as generally reliable data; if the score is lower than the second threshold, it is determined as unreliable data.
9. The intelligent integration system for multi-source heterogeneous data of territorial spatial planning based on semantic analysis according to claim 1, characterized in that Based on the unreliable data, calibrate the national territorial space planning data, specifically including: For the data marked as unreliable, call the context information in the semantic relationship graph, combine the knowledge graph reasoning mechanism, and locate the most likely semantic attribution category; use the multi-modal CLIP model to re-encode the national territorial space planning data and compare and match it with the benchmark vector in the standard semantic framework; adopt a transfer learning strategy to transfer the semantic mapping rules in the reliable data to the unreliable data to complete the semantic label correction; re-enter the calibrated data into the system for verification until the reliability evaluation standard is met.
Citation Information
Patent Citations
Multi-modal mixed teaching resource construction method and system based on cognitive neuroscience
CN117196908A
Knowledge graph driven semantic governance scheme dynamic generation method and system
CN118114758A
Visual language model-based few-sample image quality evaluation method and device
CN119316586A
Intelligent analysis system based on multi-modal medical data
CN119649977A
Unified space mapping-based agricultural multi-modal question and answer model and construction method
CN119961417A
Cited By
Land space layout adjustment decision-making method based on planning conflict identification
CN120654902A
Natural reserve intelligent checking method and system based on multi-source data driving
CN121009324A
Cross-platform-based enterprise multi-source data management system and method
CN121051128A
A cross-platform-based enterprise multi-source data management system and method
CN121051128B
Environmental protection planning dynamic generation method and system based on human-computer interaction and iterative feedback
CN121169011A