Self-adaptive semantic recognition quality inspection system and method for basic geographic update data

Through an adaptive semantic recognition and quality inspection system for basic geographic update data, the problems of quality differences in multi-source geographic data and low efficiency of traditional quality inspection methods are solved, efficient and accurate data screening, preprocessing, semantic understanding, quality inspection and correction are achieved, and the quality of geographical data and the reliability of the system are improved.

CN120179634APending Publication Date: 2025-06-20HENAN PROVINCE INST OF METROLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510189780.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively solve the difference in timeliness, accuracy and integrity of multi-source geographic data, resulting in deviations in geographical analysis and application. Traditional quality inspection methods are inefficient and prone to errors, making it difficult to deal with complex and changeable data quality problems.

Method used

Provides an adaptive semantic recognition and quality inspection system for basic geographic update data, including data acquisition module, data preprocessing module, semantic recognition module, quality inspection rule matching module, result feedback and correction module, and data update and maintenance module. The system realizes efficient data screening, preprocessing, semantic understanding, quality inspection and correction through technical means such as dynamic data source evaluation, improved wavelet transform noise reduction algorithm, geographic semantic understanding model that integrates BiLSTM and attention mechanism, multiple quality inspection rules and algorithms, automatic correction algorithm based on rule reasoning and time series prediction model.

Benefits of technology

Ensure data quality through dynamic data source evaluation, improve data preprocessing algorithms to improve data processing efficiency and accuracy, adaptive semantic recognition model accurately extracts geographic data semantic information, multiple quality inspection rules and algorithms, and automatic correction and trend analysis functions, improve the accuracy of geographic data and the reliability of the system, and reduce the cost and error rate of artificial quality inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure QLYQS_1
    Figure QLYQS_1
  • Figure QLYQS_2
    Figure QLYQS_2
  • Figure QLYQS_3
    Figure QLYQS_3
Patent Text Reader

Abstract

The invention discloses a self-adaptive semantic recognition quality inspection system and method for basic geographic update data, and particularly relates to the technical field of geographic information. The system comprises a data acquisition module, a preprocessing module, a semantic recognition module, a quality inspection rule matching module, a result feedback and correction module and a data updating and maintaining module. The data acquisition module screens high-quality data sources according to a dynamic evaluation mechanism; the preprocessing module is used for data denoising, format conversion and coordinate unification; the semantic recognition module extracts semantic information and features by using the fusion model and ResNet; the quality inspection rule matching module uses multiple algorithms to inspect data quality and dynamically update rules; the result feedback and correction module generates a report, corrects data and predicts a quality trend; and the data updating and maintaining module periodically updates data, rules and models. According to the system and the method, geographic data acquisition quality, processing efficiency and accuracy are improved, semantic and quality inspection data can be accurately identified, the system can be optimized according to feedback, and reliable data support is provided for geographic information application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of geographic information technology, and more specifically, to an adaptive semantic recognition quality inspection system and method for basic geographic updated data. Background Art

[0002] At present, with the wide application of the Geographic Information System (GIS), as the key support for many fields, the quality of basic geographic data is crucial. With the development of technology, the acquisition channels of geographic data are becoming increasingly rich, such as satellite remote sensing, unmanned aerial vehicle mapping, ground sensor networks, etc., and the data update speed is accelerating continuously. However, this also brings many challenges.

[0003] On the one hand, there are huge differences in timeliness, accuracy, and integrity among multi-source data. For example, although some satellite remote sensing images have a wide coverage range, their resolution is limited and it is difficult to meet the requirements of high-precision data for urban fine planning; ground measurement data has high accuracy, but the acquisition range is limited and the update is not timely. If the unfiltered data is directly used, it will lead to deviations in subsequent geographic analysis and applications. For example, in urban planning, due to insufficient data accuracy, the road planning does not match the actual terrain.

[0004] On the other hand, traditional quality inspection means are difficult to meet the quality inspection requirements of modern geographic data. Traditional methods are mostly based on simple rules and manual inspections, with low efficiency and easy to make mistakes. Facing a large amount of complex geographic data, manual inspection is time-consuming and laborious, and is greatly affected by subjective factors, making it difficult to ensure accuracy. At the same time, simple rules are difficult to handle complex and changeable data quality problems, such as complex spatial topological relationship errors and semantic inconsistency problems, which cannot be discovered and processed in time, affecting the reliability and practicality of the geographic information system.

[0005] Therefore, an adaptive semantic recognition quality inspection system and method for basic geographic updated data are proposed. Summary of the Invention

[0006] In order to overcome the above-mentioned defects of the prior art, the present invention provides an adaptive semantic recognition quality inspection system and method for basic geographic updated data to solve the problems raised in the above background art.

[0007] To achieve the above object, the present invention provides the following technical solution: An adaptive semantic recognition quality inspection system for basic geographic updated data, comprising:

[0008] A data acquisition module, which is used to obtain basic geographic updated data from multiple data sources, including but not limited to satellite remote sensing images, aerial photography data, and ground measurement data, and a dynamic data source evaluation mechanism is provided in the data acquisition module, according to the timeliness weight W t 、precision weight W p and integrity weight W cCalculate the comprehensive score S of the data source: S = W t ×T + W p ×P + W c ×C, where T is the timeliness score, P is the precision score, and C is the integrity score, used to automatically screen data sources with a comprehensive score higher than the preset threshold for data collection;

[0009] Data preprocessing module: Use an improved wavelet transform denoising algorithm to denoise the collected data, remove the noise in the data while retaining the key features, perform data format conversion in combination with a feature-matching-based format conversion algorithm, and use a least-squares-based coordinate conversion algorithm to unify the coordinate system;

[0010] Semantic recognition module: Construct a geographical semantic understanding model based on the fusion of a bidirectional long short-term memory network (BiLSTM) and an attention mechanism;

[0011] Quality inspection rule matching module: Used to preset quality inspection rules and standards, covering but not limited to spatial topological relationships, attribute integrity, and data consistency, and use a connectivity detection method based on the Dijkstra algorithm to check the connectivity of the road network, and use a polygon area calculation and boundary point tracking algorithm to check the closure of the building boundary;

[0012] Result feedback and correction module: Used to generate a detailed quality inspection report for the quality problems found during the quality inspection process, clearly indicating the problem type, location, and related descriptions;

[0013] Data update and maintenance module: Used to regularly update and maintain the geographical data, quality inspection rules, and semantic recognition models in the system.

[0014] Preferably, the forward and backward propagation formulas of the BiLSTM are respectively:

[0015] Forward:

[0016]

[0017]

[0018] Backward:

[0019]

[0020] Output:

[0021]

[0022] The weight formula of the attention mechanism is where e t = v T tanh(W1h t+(W2s), where v is the weight vector, W1 and W2 are weight matrices, and s is the context vector, which is used to make the model more accurately focus on key semantic information through the attention mechanism, including but not limited to understanding and extracting text descriptions, place names, and semantic information of feature attributes in geographical data;

[0023] The features of geographical data are learned and classified using a deep residual network (ResNet). The residual block formula of ResNet is:

[0024] y = F(x, {W i ) + x

[0025] where x is the input, F(x, {W i ) is the residual function, and {W i} is the set of weights used to identify different geographical features. In model training, transfer learning technology is adopted, and the loss function is:

[0026] L = L task + λL pre

[0027] where L task is the loss of the current task, L pre is the loss of the pre-trained model, and λ is the balance coefficient, which is used to accelerate the convergence and optimization of the model by leveraging the existing large-scale pre-trained geographical data model

[0028] Preferably, the distance update formula of the Dijkstra algorithm in the quality inspection rule matching module is:

[0029] dist[v] = min(dist[v], dist[u] + w(u, v)

[0030] where dist[v] is the shortest distance from the source point to vertex v, u is the vertex with the determined shortest path, and w(u, v) is the weight of the edge (u, v);

[0031] The formula for calculating the area of a polygon is:

[0032]

[0033] where (x i , y i ) are the vertex coordinates of the polygon, n is the number of vertices, and the boundary point tracking adopts the Moore neighborhood tracking algorithm; by setting the threshold [V min , V max of the attribute value range, it is judged whether the feature attribute value V is within a reasonable range, that is, it is judged whether V min ≤ V ≤ V max holds;

[0034] This module also has the function of dynamic rule update. According to the actual quality inspection results and user feedback, the rule update probability formula is adopted:

[0035]

[0036] where N error is the number of quality inspection errors, and N total is the total number of quality inspections. When P update >α, the quality inspection rules are automatically adjusted and improved, where α is the preset update probability threshold.

[0037] Preferably, when the result feedback and correction module runs, for the data with incorrect attribute values, an automatic correction algorithm based on rule reasoning is used for correction, and the rule reasoning adopts the Bayesian reasoning formula:

[0038]

[0039] where H is the hypothesis, E is the evidence, P(H|E) is the posterior probability, P(E|H) is the likelihood probability, P(H) is the prior probability, and P(E) is the probability of the evidence;

[0040] For the data with incorrect spatial topological relationships, a visual error prompt is generated to guide the data processing personnel to make manual corrections; this module also adds a quality trend analysis function. By analyzing the historical quality inspection data, the time series prediction model y t =β0 + β1yt -1 +β2yt -2 +…+β p y t-p +∈ t is adopted, where y t is the data quality index at time t, β i is the model parameter, p is the lag order, and ∈ t is the random error term, which is used to predict the change trend of data quality and take preventive measures in advance

[0041] Preferably, during the training process of the geographical semantic understanding model that fuses BiLSTM and attention mechanism in the semantic recognition module, the adaptive learning rate adjustment algorithm Adagrad is adopted, and the learning rate update formula is:

[0042]

[0043] where η0 is the initial learning rate, G t is the sum of the squared gradients at time t, and ∈ is a small constant to prevent division by zero.

[0044] Preferably, during the training process of the deep residual network (ResNet) in the semantic recognition module, the stochastic gradient descent (SGD) combined with the momentum optimization algorithm is adopted, and the parameter update formula is:

[0045] v t+1 = γv t + ηΔL(θ t )

[0046] θ t+1 = θ t - v t+1

[0047] where v t is the momentum at time t, γ is the momentum coefficient, η is the learning rate, ΔL(θ t ) is the gradient at time t, and θ t is the parameter at time t, which is used to accelerate the convergence of the model, prevent the model from falling into local optima, and introduce the L2 regularization technique. The regularization term formula is:

[0048]

[0049] where τ is the regularization coefficient and θ i is the model parameter, which is used to reduce the overfitting phenomenon of the model.

[0050] Preferably, in the connectivity detection method based on the Dijkstra algorithm in the quality inspection rule matching module, the parallel computing technology is adopted, and the number of parallel computing threads N thread is dynamically adjusted according to the data scale N data , and the calculation formula is where M is the amount of data processed by each thread, and N max is the maximum number of threads, which is used to improve the detection efficiency.

[0051] Preferably, in the quality inspection rule matching module, the polygon area calculation and boundary point tracking algorithm are combined with the topology optimization technology, and the objective function of the topology optimization is

[0052]

[0053] where x is the design variable, c i is the cost coefficient, and the constraint condition is g j (x) ≤ 0 (j = 1, 2,..., m), which is used to more accurately process complex building boundary situations.

[0054] A quality inspection method for an adaptive semantic recognition quality inspection system for basic geographic update data according to the above-mentioned claims includes the following steps:

[0055] Data collection step: Obtain basic geographic update data from multiple data sources; during the collection process, intelligently screen the highest-quality data according to the preset data source priority and data quality assessment indicators.

[0056] Data preprocessing step: Use a denoising algorithm based on wavelet transform to denoise the collected data, combine a format conversion algorithm based on feature matching for data format conversion, and adopt a coordinate transformation algorithm based on the least squares method to unify the coordinate system; at the same time, use a data augmentation algorithm to expand the data volume.

[0057] Semantic recognition step: Construct a geographic semantic understanding model based on the fusion of bidirectional long short-term memory network (BiLSTM) and attention mechanism to understand and extract semantic information such as text descriptions, place names, and feature attributes in geographic data; use a deep residual network (ResNet) to learn and classify the features of geographic data to identify different geographic features; adopt transfer learning technology during model training.

[0058] Quality inspection rule matching step: Preset quality inspection rules and standards covering multiple aspects such as spatial topological relationships, attribute integrity, and data consistency. Use a connectivity detection method based on the Dijkstra algorithm to check the connectivity of the road network, use a polygon area calculation and boundary point tracking algorithm to check the closure of building boundaries, and judge whether the feature attribute values are within a reasonable range by setting the threshold of the attribute value range; dynamically update the quality inspection rules according to the actual quality inspection results and user feedback.

[0059] Result feedback and correction step: For quality problems found during the quality inspection process, generate a detailed quality inspection report clearly indicating the problem type, location, and relevant descriptions; for data with incorrect attribute values, use an automatic correction algorithm based on rule reasoning for correction; for data with incorrect spatial topological relationships, generate a visual error prompt to guide data processing personnel for manual correction; at the same time, analyze historical quality inspection data to predict the change trend of data quality.

[0060] Data update and maintenance step: Regularly update and maintain the geographic data, quality inspection rules, and semantic recognition models in the system; monitor the changes in data sources in real time and update the collected data in a timely manner.

[0061] The technical effects and advantages of the present invention:

[0062] (1) Through a dynamic data source evaluation mechanism, calculate the comprehensive score of data sources based on the weights of timeliness, accuracy, and integrity, screen high-quality data sources, ensure the quality of the collected data, and improve the reliability of subsequent analysis. For example, in an urban geographic data update project, it can accurately select, from numerous data sources, such as the latest high-resolution satellite remote sensing images and precise ground measurement data, providing a reliable basis for subsequent applications such as urban planning.

[0063] (2) Effectively remove noise and retain key features through an improved wavelet transform denoising algorithm; implement data compatibility for different formats based on a feature-matching format conversion algorithm; use the least squares method for coordinate transformation to unify the coordinate system, improve the efficiency and accuracy of data processing, and reduce the impact of data errors on subsequent work. For example, when processing terrain data, the terrain features are clearer after denoising, facilitating terrain analysis.

[0064] (3) A geographical semantic understanding model that combines BiLSTM and attention mechanisms, combined with ResNet and transfer learning techniques, can accurately extract semantic information and features of geographical data, identify different geographical elements, and has high training efficiency and strong generalization ability. In the processing of urban geographical data, various types of buildings, roads, etc. can be accurately identified and their attribute semantics can be understood.

[0065] (4) The quality inspection rule matching module covers a variety of quality inspection rules, uses multiple algorithms to check spatial topological relationships, attribute integrity, and data consistency, and also has a function of dynamically updating rules, which can timely detect and adapt to changes in data quality, ensuring the accuracy of geographical data. For example, it can accurately detect the connectivity of the road network and the closure of building boundaries, and timely discover data problems.

[0066] (5) Generate a detailed quality inspection report through the result feedback and correction module to facilitate problem location; use Bayesian inference to automatically correct data with incorrect attribute values, and visually prompt to manually correct data with incorrect topological relationships; analyze the quality trend through a time series prediction model for early prevention, improving data quality and system usability. For example, by predicting the downward trend of data quality, optimize the acquisition link in advance.

[0067] (6) Regularly update geographical data, quality inspection rules, and semantic recognition models through the data update and maintenance module, monitor changes in data sources in real time and update in a timely manner, and adopt an incremental update strategy to save resources, ensuring the stable operation of the system and the timeliness of data. For example, update data according to new data sources, and optimize models and rules to adapt to new requirements. Specific implementation manners

[0068] Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0069] An adaptive semantic recognition quality inspection system for basic geographical update data provided by the present invention includes:

[0070] I. Data acquisition module

[0071] (1) Data acquisition

[0072] After the system starts, the data acquisition module begins to work, obtaining basic geographic update data from various data sources such as satellite remote sensing images, aerial photography data, and ground measurement data. For example, in an urban geographic data update project, large-scale topographic and geomorphic information is obtained from satellite remote sensing images, and accurate location information of urban roads and buildings is obtained from ground measurement data.

[0073] (2) Evaluation and Screening of Dynamic Data Sources

[0074] Evaluation Calculation: The data acquisition module determines the timeliness weight W t , accuracy weight W p and integrity weight W c for each data source according to the dynamic data source evaluation mechanism, and calculates the comprehensive score S of the data source as S = W t ×T + W p ×P + W c ×C. Specifically, in a certain project, W t is set to 0.3, W p is set to 0.4, and W c is set to 0.3 according to actual needs. Then, the timeliness score T, accuracy score P, and integrity score C are calculated for each data source. Taking a certain satellite remote sensing image data source as an example, if its acquisition time is recent, the timeliness score T = 0.8; after detecting that its data accuracy meets the project requirements, the accuracy score P = 0.7; the data coverage is complete, and the integrity score C = 0.8. The comprehensive score of the data source is calculated through the formula S = W t ×T + W p ×P + W c ×C, that is, S = 0.3×0.8 + 0.4×0.7 + 0.3×0.8 = 0.76.

[0075] Screening Decision: Preset a comprehensive score threshold for the data source, such as S threshold = 0.7. Compare the comprehensive scores of each calculated data source with the threshold, and automatically screen out the data sources with comprehensive scores higher than the preset threshold for data acquisition. In the above example, the comprehensive score of this satellite remote sensing image data source is 0.76 > 0.7, so it is selected for data acquisition.

[0076] II. Data Preprocessing Module

[0077] The collected data is denoised using an improved wavelet transform denoising algorithm to remove the noise in the data while retaining the key features. Then, the denoised data is obtained by performing wavelet reconstruction on the processed wavelet coefficients. Specifically, when processing the terrain data with noise, after denoising, the contour and detail features of the terrain become clearer, providing a more accurate data basis for subsequent analysis. The data format is converted by combining a format conversion algorithm based on feature matching. According to the pre-set conversion rules, the data is converted from one format to another. Specifically, the vector format road data is converted to the raster format for subsequent processing under a unified data format. The coordinate system is unified by using a coordinate conversion algorithm based on the least squares method, so that the building coordinates recorded in different coordinate systems in different data sources are uniformly converted to the coordinate system required by the project, ensuring the consistency of all data in terms of spatial position and facilitating subsequent spatial analysis and processing.

[0078] III. Semantic Recognition Module

[0079] (I) Model Construction and Training

[0080] Construction of BiLSTM and Attention Mechanism Model: A geographical semantic understanding model based on the fusion of bidirectional long short-term memory network (BiLSTM) and attention mechanism is constructed. In the model structure design, the BiLSTM layer is responsible for learning the time series features of geographical data and capturing the context information of the data. The attention mechanism is used to enhance the model's ability to focus on key semantic information.

[0081] Model Training: Prepare training data, including a large number of annotated geographical data samples, such as text data containing place names and descriptions of geographical features. During the training process, transfer learning technology is adopted, and the parameters of the pre-trained model with existing large-scale geographical data are used as the initial values to accelerate the convergence speed of the model. The loss function is L = L task + λL pre , where L task is the loss of the current task, L pre is the loss of the pre-trained model, and λ is the balance coefficient, whose value is adjusted according to experiments to optimize the training effect.

[0082] Learning Rate Adjustment: During the training of the geographical semantic understanding model integrating BiLSTM and attention mechanism, the adaptive learning rate adjustment algorithm Adagrad is adopted. The initial learning rate is set to η0 = 0.001. During the training process, the learning rate is updated according to the formula . Where G t is the sum of the squared gradients at time t, and ∈ is a small constant to prevent division by zero, set to 1e-8. As the training progresses, the learning rate will be automatically adjusted according to the change of the gradient, improving the training efficiency.

[0083] ResNet Model Training: Use the Deep Residual Network (ResNet) to learn and classify the features of geographical data. The residual block formula of ResNet is y = F(x, {W i}) + x, where x is the input, F(x, {W i}) is the residual function, and {W i} is the weight set. When training ResNet, the Stochastic Gradient Descent (SGD) combined with the momentum optimization algorithm is adopted, with the momentum coefficient γ = 0.9 and the learning rate η = 0.001. The parameter update formula is v t+1 = γv t + ηΔL(θ t ), θ t+1 = θ t - v t+1 . At the same time, the L2 regularization technique is introduced, with the regularization coefficient τ = 0.0001, and the regularization term formula is:

[0084]

[0085] to reduce the overfitting phenomenon of the model.

[0086] (II) Semantic Recognition and Feature Extraction

[0087] Data Input and Processing: The preprocessed data is input into the semantic recognition module. For data such as text descriptions and place names, preprocessing operations such as word segmentation and encoding are first performed to convert it into an input format acceptable to the model. For image data, such as satellite remote sensing images, operations such as cropping and normalization are performed.

[0088] Semantic Understanding and Feature Extraction: The data is input into a model that fuses BiLSTM and the attention mechanism. The model, based on the knowledge learned during training, understands and extracts semantic information such as text descriptions, place names, and feature attributes of geographical data. The attention mechanism is through the formula where e t = v T tan h(W1h t + W2s), v is the weight vector, W1 and W2 are the weight matrices, and s is the context vector. Through the attention weights, the model can focus on key semantic information and improve the accuracy of semantic understanding. At the same time, the ResNet model learns and classifies the image features of geographical data, identifying different geographical elements such as buildings, roads, and rivers. Specifically, when processing urban geographical data, it can accurately identify different types of buildings and extract their feature information.

[0089] IV. Quality Inspection Rule Matching Module

[0090] (I) Quality Inspection Rule Setting

[0091] Preset quality inspection rules and standards covering multiple aspects such as spatial topological relationships, attribute integrity, and data consistency. For example, in terms of spatial topological relationships, it is stipulated that the road network should be connected and the building boundaries should be closed; in terms of attribute integrity, it is required that each feature must have the specified attribute information, such as the height and usage of buildings cannot be missing; in terms of data consistency, it is required that the attribute values of the same feature in different data sources should be consistent.

[0092] (2) Connectivity detection

[0093] Construct a graph structure: Use a connectivity detection method based on the Dijkstra algorithm to check the connectivity of the road network. Abstract the road network as a graph structure, with road nodes as the vertices of the graph and road connection relationships as the edges. Assign a weight w(u,v) to each edge, and the weight can represent information such as the length and traffic capacity of the road.

[0094] Execute the Dijkstra algorithm: Start from the selected source point, initialize the distance dist[v] from the source point to each vertex to infinity, and the distance from the source point to itself is 0. Then, perform iterative updates according to the distance update formula dist[v]=min(dist[v],dist[u]+w(u,v)) of the Dijkstra algorithm. In each iteration, select the vertex u that is closest to the source point and whose shortest path has not been determined, and update the distance of its adjacent vertex v. After multiple iterations, until the shortest distances of all vertices are determined.

[0095] Connectivity judgment: Judge whether the road network is connected according to the calculation results. If the shortest distances of all vertices are not infinity, it means the road network is connected; otherwise, there are disconnected parts in the road network. For example, in the quality inspection of an urban road network, it is found through this algorithm that there are breakpoints in a certain area of the road network, resulting in some areas being unable to be directly reached by road, which requires further inspection and repair.

[0096] (3) Building boundary inspection

[0097] Polygon area calculation: Based on polygon area calculation and boundary point tracking algorithm, check the closure of the building boundary. For the polygon representing the building boundary, calculate its area according to the formula

[0098]

[0099] where (x i ,y i ) are the vertex coordinates of the polygon and n is the number of vertices. If the calculated area is 0 or close to 0, there may be a situation where the boundary is not closed.

[0100] Boundary point tracking: The Moore neighborhood tracking algorithm is used to track the boundary points of buildings. Starting from a starting point of the polygon, adjacent boundary points are sequentially tracked in a certain order (such as clockwise or counterclockwise), and it is checked whether all boundary points can be traversed completely and return to the starting point. If it is impossible to continue tracking or return to the starting point during the tracking process, it indicates that the building boundary is not closed.

[0101] Attribute value judgment: By setting the threshold of the attribute value range [V min , V max to judge whether the feature attribute value V is within a reasonable range. For example, for the height attribute of buildings, the reasonable range is set as [1, 100] (unit: meter). If the height attribute value of a certain building is not within this range, it is determined that the attribute value is unreasonable.

[0102] (IV) Rule dynamic update

[0103] Error statistics: During the quality inspection process, record the number of quality inspection errors N error and the total number of quality inspections N total . For example, in a quality inspection of urban geographic data, a total of 1000 features were inspected, and it was found that 50 of them had various quality problems. Then N error = 50, N total = 1000.

[0104] Rule update judgment: Calculate the rule update probability according to the rule update probability formula . Preset an update probability threshold α = 0.05. When P update > α, it indicates that the current quality inspection rules may be imperfect and need to be updated. In the above example, it just reaches the threshold, and the system triggers the rule update mechanism.

[0105] Rule update execution: The system automatically adjusts and improves the quality inspection rules according to the actual quality inspection results and user feedback. For example, if it is found that the height attribute values of a large number of buildings are misjudged as unreasonable, it may be that the threshold setting is unreasonable. The system will readjust the threshold range of the height attribute according to the actual situation, or add other relevant judgment conditions to improve the accuracy and adaptability of the quality inspection rules.

[0106] V. Result feedback and correction module

[0107] (I) Quality inspection report generation

[0108] Problem Detection and Recording: After the quality inspection rule matching module completes the quality inspection, the result feedback and correction module starts to work. For the quality problems found during the quality inspection, the system details record the problem type, location, and relevant description. For example, when inspecting the geographical data of a certain area, it is found that the boundary of a certain building is not closed. The recorded problem type is "building boundary error", the location is the coordinate position of the building in the geographical data, and the relevant description is "a break point in the building boundary is detected through polygon area calculation and boundary point tracking algorithm".

[0109] Report Generation: According to the recorded problem information, a detailed quality inspection report is generated. The quality inspection report is presented in a clear and easy-to-understand format, including the summary statistics of the problems, the specific details of each problem, etc. For example, the quality inspection report may show that a total of 10 quality problems are found in this quality inspection, including 3 problems with road network connectivity, 5 problems with building boundaries, and 2 problems with unreasonable attribute values, and the specific location and detailed description of each problem are listed separately.

[0110] (2) Data Correction

[0111] Attribute Value Correction: For the data with incorrect attribute values, an automatic correction algorithm based on rule reasoning is used for correction. Rule reasoning uses the Bayesian reasoning formula Specifically, when judging the usage attribute of a certain building, given the environmental information (evidence E) around the building, the prior probability P(H) of different usage buildings appearing in this area, and the likelihood probability P(E|H) of the current surrounding environmental information appearing under different usages. The posterior probability P(H|E) is calculated through the Bayesian reasoning formula, so as to infer the most likely usage attribute of the building and correct the incorrect usage attribute.

[0112] Spatial Topological Relationship Correction: For the data with incorrect spatial topological relationships, such as problems with unconnected road networks or unclosed building boundaries, a visual error prompt is generated. On the geographical information visualization interface, the incorrect topological relationship part is highlighted in a prominent color or mark to guide the data processing personnel to make manual corrections. For example, the unconnected road sections are marked with red lines, and the unclosed building boundaries are marked with flashing yellow lines, which is convenient for the data processing personnel to quickly locate and fix the problems.

[0113] (3) Quality Trend Analysis

[0114] Data Collection and Sorting: Collect historical quality inspection data, including information such as the time of each quality inspection, the problem types found, and the quantity. These data are sorted out to extract the data quality index y t , for example, the number of problems found in each quality inspection can be used as the data quality index.

[0115] Model Training and Prediction: Use a time series prediction model y t = β0 + β1y t-1 + β2y t-2 + … + β p yt -p + ∈ t Conduct data quality trend analysis. Train the model with historical quality inspection data to determine the model parameters β i . Specifically, train with the quality inspection data of the past 12 months to obtain the model parameters. Then, use the trained model to predict the future data quality change trend. Assume that the prediction result shows a downward trend in data quality in the coming months, and take preventive measures in advance, such as strengthening quality control in the data collection process or optimizing the quality inspection rules.

[0116] VI. Data Update and Maintenance Module

[0117] Regularly update and maintain geographical data, quality inspection rules, and semantic recognition models. Monitor the changes in data sources in real time and update the collected data in a timely manner when data changes. When updating geographical data, adopt an incremental update strategy to reduce update time and resource consumption. Thereby ensuring the accuracy and timeliness of the data in the system, enabling the quality inspection rules and semantic recognition models to adapt to new data characteristics and requirements, ensuring the continuous and stable operation of the entire quality inspection system, and providing long-term guarantee for the quality of geographical data.

[0118] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. An adaptive semantic recognition quality inspection system for basic geographic update data, characterized in that: include: The data acquisition module is used to obtain basic geographic update data from a variety of data sources, including but not limited to satellite remote sensing images, aerial photography data, and ground measurement data. The data acquisition module is equipped with a dynamic data source evaluation mechanism, which is based on the timeliness weight W of the data. t , precision weight W p and the complete weight W c Calculate the comprehensive score of data source S = W t ×T+W p ×P+W c ×C, where T is the timeliness score, P is the precision score, and C is the completeness score, which is used to automatically screen data sources with comprehensive scores higher than the preset threshold for data collection; The data preprocessing module uses an improved wavelet transform denoising algorithm to denoise the collected data, remove the noise in the data, and retain the key features. It combines the format conversion algorithm based on feature matching to convert the data format, and uses the coordinate conversion algorithm based on the least squares method to achieve coordinate system unification. Semantic recognition module, which builds a geographic semantic understanding model based on the fusion of bidirectional long short-term memory network (BiLSTM) and attention mechanism; The quality inspection rule matching module is used to pre-set quality inspection rules and standards, covering but not limited to spatial topological relationships, attribute integrity, and data consistency. It also uses the connectivity detection method based on the Dijkstra algorithm to check the connectivity of the road network, and checks the closure of the building boundary based on polygon area calculation and boundary point tracking algorithm; The result feedback and correction module is used to generate a detailed quality inspection report for quality problems found during the quality inspection process, clearly indicating the problem type, location and related description; The data update and maintenance module is used to regularly update and maintain the geographic data, quality inspection rules and semantic recognition models in the system.

2. The adaptive semantic recognition quality inspection system for basic geographic update data according to claim 1 is characterized in that: The forward and backward propagation formulas of the BiLSTM are: Forward: Backward: Output: The weight formula of the attention mechanism is where e t =v T tanh(W1h t +W2s), v is the weight vector, W1 and W2 are weight matrices, and s is the context vector, which is used to enable the model to focus more accurately on key semantic information through the attention mechanism, including but not limited to understanding and extracting text descriptions, place names, and semantic information of geographical attributes in geographic data; The deep residual network (ResNet) is used to learn and classify the features of geographic data. The residual block formula of ResNet is: y=F(x,{W i })+x Among them, x is the input, F(x,{W i }) is the residual function, {W i } is a weight set used to identify different geographical elements; in model training, transfer learning technology is used, and the loss function is: L=L task +λL pre Among them, L task is the loss of the current task, L pre is the loss of the pre-trained model, and λ is the balancing coefficient, which is used to accelerate the convergence and optimization of the model with the help of the existing large-scale geographic data pre-trained model.

3. The adaptive semantic recognition quality inspection system for basic geographic update data according to claim 1, characterized in that: The distance update formula of the Dijkstra algorithm in the quality inspection rule matching module is: dist[v]=min(dist[v],dist[u]+w(u,v) Where dist[v] is the shortest distance from the source point to vertex v, u is the vertex of the determined shortest path, and w(u,v) is the weight of the edge (u,v); The formula for calculating the area of ​​a polygon is: Where (x i ,y i ) is the vertex coordinate of the polygon, n is the number of vertices, and the boundary point tracking adopts Moore's neighborhood tracking algorithm; by setting the threshold value of the attribute value range [V min ,V max ] Determine whether the feature attribute value V is within a reasonable range, that is, determine whether V min ≤V≤V max whether it is established; This module also has the function of dynamically updating rules. According to the actual quality inspection results and user feedback, the rule update probability formula is adopted: Where N error is the number of quality inspection errors, N total is the total number of quality inspections, when P update >α, the quality inspection rules are automatically adjusted and improved, where α is the preset update probability threshold.

4. The adaptive semantic recognition quality inspection system for basic geographic update data according to claim 1, characterized in that: The result feedback and correction module operates by using an automatic correction algorithm based on rule reasoning to correct data with incorrect attribute values. The rule reasoning adopts the Bayesian reasoning formula: Where H is the hypothesis, E is the evidence, P(H|E) is the posterior probability, P(E|H) is the likelihood probability, P(H) is the prior probability, and P(E) is the probability of the evidence; For data with incorrect spatial topological relationships, a visual error prompt is generated to guide data processing personnel to make manual corrections; this module also adds a quality trend analysis function, which uses the time series prediction model y through the analysis of historical quality inspection data. t =β0+β1y t-1 +β2y t-2 +…+β p y t-p +∈ t , where y t is the data quality index at time t, β i is the model parameter, p is the lag order, ∈ t It is a random error term, which is used to predict the changing trend of data quality and take preventive measures in advance.

5. The adaptive semantic recognition quality inspection system for basic geographic update data according to claim 1 is characterized in that: In the training process of the geographic semantic understanding model that integrates BiLSTM and attention mechanism in the semantic recognition module, the adaptive learning rate adjustment algorithm Adagrad is adopted, and the learning rate update formula is: Where η0 is the initial learning rate, G t is the sum of squared gradients at time t, and ∈ is a small constant to prevent division by zero.

6. The adaptive semantic recognition quality inspection system for basic geographic update data according to claim 2 is characterized in that: In the training process of the deep residual network (ResNet) in the semantic recognition module, the stochastic gradient descent (SGD) combined with the momentum optimization algorithm is used, and the parameter update formula is: v t+1 =γv t +ηΔL(θ t ) i t+1 =θ t -v t+1 Among them, v t is the momentum at time t, γ is the momentum coefficient, η is the learning rate, ΔL(θ t ) is the gradient at time t, θ t is the parameter at time t, which is used to accelerate the convergence of the model, prevent the model from falling into the local optimum, and introduce the L2 regularization technology. The regularization term formula is: Among them, τ is the regularization coefficient, θ i is a model parameter used to reduce the overfitting phenomenon of the model.

7. The adaptive semantic recognition quality inspection system for basic geographic update data according to claim 1 is characterized in that: The connectivity detection method based on Dijkstra algorithm in the quality inspection rule matching module adopts parallel computing technology, and the number of parallel computing threads N thread According to the data size N data Dynamic adjustment, the calculation formula is: Where M is the amount of data processed by each thread, N max The maximum number of threads, used to improve detection efficiency.

8. The adaptive semantic recognition quality inspection system for basic geographic update data according to claim 1, characterized in that: The quality inspection rule matching module is based on polygon area calculation and boundary point tracking algorithm combined with topology optimization technology. The objective function of topology optimization is: Where x is the design variable, c i is the cost coefficient, and the constraint g j (x)≤0(j=1,2,…,m), which is used to handle complex building boundary situations more accurately.

9. A quality inspection method for an adaptive semantic recognition quality inspection system for basic geographic update data according to any one of claims 1 to 8, characterized in that: The following steps are involved: Data collection step, obtaining basic geographic update data from various data sources; During the collection process, the best quality data is intelligently selected based on the preset data source priority and data quality assessment indicators; In the data preprocessing step, the collected data is denoised using a denoising algorithm based on wavelet transform, and the data format is converted using a format conversion algorithm based on feature matching. The coordinate system is unified using a coordinate conversion algorithm based on the least squares method. At the same time, the data enhancement algorithm is used to expand the data volume. In the semantic recognition step, a geographic semantic understanding model based on the fusion of BiLSTM and attention mechanism is constructed to understand and extract semantic information such as text descriptions, place names, and attributes of geographical objects in geographic data; the deep residual network (ResNet) is used to learn and classify the features of geographic data and identify different geographical elements; transfer learning technology is used in model training; Quality inspection rule matching step: pre-set quality inspection rules and standards, covering multiple aspects such as spatial topological relationship, attribute integrity, data consistency, etc., using the connectivity detection method based on the Dijkstra algorithm to check the connectivity of the road network, based on polygon area calculation and boundary point tracking algorithm to check the closure of the building boundary, and by setting the threshold of the attribute value range to determine whether the attribute value of the object is within a reasonable range; Dynamically update quality inspection rules based on actual quality inspection results and user feedback; Result feedback and correction steps: For quality problems found during the quality inspection process, a detailed quality inspection report is generated, clearly indicating the problem type, location and related description; for data with incorrect attribute values, an automatic correction algorithm based on rule reasoning is used to correct them; for data with incorrect spatial topological relationships, a visual error prompt is generated to guide data processing personnel to make manual corrections; at the same time, historical quality inspection data is analyzed to predict the changing trend of data quality; Data update and maintenance steps: regularly update and maintain the geographic data, quality inspection rules and semantic recognition models in the system; monitor changes in data sources in real time and update collected data in a timely manner.