Method for detecting defects of two-dimensional vector geographic database and computer readable storage medium

By calculating the quality risk index of sample units and using a random sampling method, high-risk sample units are selected for testing, which solves the problem of missed defects in two-dimensional vector geographic databases and achieves more efficient and reliable detection results.

CN120873101BActive Publication Date: 2025-11-25SICHUAN SURVEYING & MAPPING PROD QUALITY SUPERVISION & INSPECTION STATION OF THE MINIST OF NATURAL RESOURCES SICHUAN SURVEYING & MAPPING PROD QUALITY SUPERVISION & INSPECTION STATION
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511396362.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-28
Publication Date
2025-11-25
Estimated Expiration
2045-09-28

AI Technical Summary

Technical Problem

Existing technologies fail to effectively identify differences in spatial location and feature distribution characteristics in two-dimensional vector geographic database defect detection, leading to missed defects and reducing detection accuracy and the reliability of inspection conclusions.

Method used

By calculating the quality risk index of sample units, high-risk sample units are screened out. Random and uniform sampling methods are used, combined with automated detection tools, to sort and detect the sample units, ensuring that high-risk samples are sampled with a high probability. The inspection is stopped when non-compliance is detected.

Benefits of technology

It improves the accuracy of defect detection and the reliability of inspection conclusions in two-dimensional vector geographic databases, reduces inspection time and cost, and ensures the objectivity and accuracy of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120873101B_ABST
    Figure CN120873101B_ABST
Patent Text Reader

Abstract

The application discloses a kind of two-dimensional vector geographic database defect detection method and computer readable storage medium, belong to electric digital data processing field.For solving the risk of reducing the defect of two-dimensional vector geographic database is missed, the defect detection method of the present application includes: screening module receives inspection batch, determines sample total number and submission number;Calculate quality risk index, obtain the first coordinate sequence of geometric center coordinate;Random number is generated and the second coordinate sequence is constructed, and then the corresponding sample unit and quality risk index are taken;Random number is generated and compared with quality risk index, to determine repeated sampling or join submission set;Submission set is in descending order according to quality risk index and detected using automated tool;Decision module encounters unqualified and refuses to receive the inspection batch.The present application is applied to database field, with the advantages of low database defect missed detection rate, high reliability of detection result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electronic digital data processing, and specifically to a method for detecting defects in a two-dimensional vector geographic database and a computer-readable storage medium. Background Technology

[0002] A database (usually referring to a relational database, the mainstream type of traditional database) is fundamentally about organizing data in tabular form and manipulating it using Structured Query Language (SQL) to simultaneously ensure data security, consistency, integrity, and scalability. It is a software system used for the storage, management, querying, and maintenance of structured data. Its core objective is to solve the problem of "how data can be stored and used efficiently, securely, and systematically," and it is the core data carrier of modern information systems (such as websites, applications, and enterprise management systems).

[0003] Two-dimensional vector geodatabases are the core data carriers of Geographic Information Systems (GIS) and are currently widely used surveying and mapping products. They are extensively applied in fields requiring precise spatial analysis, such as urban planning and traffic management. Two-dimensional vector geodatabases also serve as the infrastructure for geospatial data management. They abstract real-world geographic elements through vector forms—vector points, vector lines, and vector polygons—describing real-world geographic elements (such as roads, buildings, and rivers) using geometric objects like points, lines, and polygons, and associating them with attribute information (such as road names and building heights). This deeply binds spatial location with attribute information and supports core capabilities such as spatial queries and topological rules, ultimately providing data support for precise decision-making in fields such as urban planning, transportation, and natural resources. Two-dimensional vector geodatabases are typically in file database formats such as MDB and GDB. Each point, line, or area is called a geographic information element, and is divided into nine major types: location basis, water network, settlement and facilities, transportation data network, pipelines, boundaries and administrative divisions, landforms, vegetation and soil, and place names.

[0004] On the one hand, two-dimensional vector geodatabases simultaneously store both "spatial information" (such as the latitude and longitude coordinates and shape range of a school) and "attribute information" (such as school name, type of school, and number of students), supporting spatial relationship operations (such as querying supermarkets within 500 meters of a road or determining whether two plots are adjacent), and ensuring spatial consistency of data (such as avoiding overlapping or disconnection of the boundaries of two adjacent plots). On the other hand, the storage methods of two-dimensional vector geodatabases are mainly divided into two categories: "vector" and "raster," with vector being the core data format of two-dimensional vector geodatabases. Furthermore, a "two-dimensional" vector geodatabase refers to a geospatial feature described only using a planar coordinate system (such as latitude and longitude or Cartesian coordinates), excluding the third dimension of "elevation (altitude)."

[0005] Common methods for constructing two-dimensional vector geographic databases include: collecting ground data using surveying instruments such as total stations, Global Positioning System Real Time Kinematic (GPS-RTK), and airborne or vehicle-mounted lidar; obtaining ground spatial distribution data through remote sensing classification and other methods using satellite remote sensing images or aerial photographs; and directly using existing digital map data and converting it to the required format for a two-dimensional vector geographic database.

[0006] However, deviations can easily occur in the entire process from "real-world geographic entities" to "digital vector representations," particularly in information transmission, processing, and storage. The deficiencies in two-dimensional vector geodatabases primarily stem from errors in data production, deviations in data processing and integration, and failures in data storage and maintenance. For example, vector data production relies on original data sources (such as remote sensing imagery, aerial photographs, and field measurement records). If the data source itself lacks precision, it will be directly transmitted to the vector data. Whether it's field data collection (such as GPS measurements or total station measurements) or indoor digitization (such as vectorization of paper maps or image interpretation), errors can be introduced due to technical limitations or human error. Furthermore, errors in coordinate systems and projection transformations, failures in constructing data topology relationships, and deviations in matching and fusing multi-source data are common and cumulative deviations in data processing and integration. Finally, damage to storage media (such as bad sectors on hard drives), transmission interruptions (such as network instability), and untimely or incorrect data updates can all lead to data corruption or incompleteness.

[0007] Due to various non-ideal factors in reality, the geospatial data in the constructed two-dimensional vector geographic database is not always accurate and reliable. Therefore, in practice, it is often necessary to conduct quality inspections on digital mapping results to detect defects in the two-dimensional vector geographic database and thus determine whether the digital mapping results are qualified.

[0008] Because two-dimensional vector geographic databases are extremely large, the computational complexity required for detection often matches the square law of geographic information elements. Therefore, it is impractical to check the accuracy of every single geospatial data point. According to relevant national standards, current technologies treat geospatial data as homogeneous, standardized products. Directly applying existing theories is relatively simple, employing random sampling or stratified random sampling for defect detection, assigning each data point the same probability of inclusion. However, current technologies fail to recognize a core flaw: ignoring the fact that different sample data have different spatial locations and different feature distribution characteristics, treating geospatial data like standard industrial products. When processing batches of products containing a small amount of highly complex, high-risk, and high-application-value data, it cannot guarantee that these key samples will receive a higher probability of inclusion, leading to easily missed defects. This missed defect phenomenon directly increases the quality risk of the batch of products and ultimately reduces the reliability of quality inspection conclusions.

[0009] Prior art 1: CN119557385B, A cloud computing-based surveying and mapping geographic information data management method and system, publication date: April 29, 2025;

[0010] Prior art 2: CN118071208A, Sampling method for quality inspection of digital line plots using spatial heterogeneity, Publication date: May 24, 2024.

[0011] Existing technology 1 addresses the problem of dynamic prediction of geographic elements by employing a spatiotemporal prediction sequence model based on the Spatiotemporal Graph Neural Network (STGCN) to predict the evolution trend of spatiotemporal attribute maps, thereby achieving dynamic prediction. Existing technology 2 calculates spatial heterogeneity based on the land cover types of point, line, and area elements using Euclidean nearest neighbor distance, gyration radius exponent, and fractal dimension exponent, respectively. Finally, it uses samples selected based on spatial heterogeneity to verify the quality of digitized line maps. This approach aims to ensure the accuracy of sample unit detection while reducing information redundancy and improving verification efficiency; it is a technical solution applicable to automated tools. In other words, neither existing technology 1 nor existing technology 2 addresses the deficiencies of the prior art proposed in this invention.

[0012] Given the short timeframe, limited funding, high cost, and massive computational demands of acceptance testing for 2D vector geodatabases, a stark contrast exists between the large volume of data to be tested within these databases. Therefore, effectively improving the accuracy of defect detection and the reliability of test results for 2D vector geodatabases is one of the key technical challenges in this field. Summary of the Invention

[0013] To alleviate or partially alleviate the above-mentioned technical problems, the solution of the present invention is as follows:

[0014] A method for detecting defects in a two-dimensional vector geographic database includes:

[0015] Step S1: The sample unit screening module receives the test batch, obtains the number N of sample units in the test batch, and configures the number of samples to be sent for testing, n.

[0016] Step S2: Calculate and obtain the quality risk index of each sample unit in the inspection batch;

[0017] Step S3: Obtain the position coordinates of the geometric center of each sample unit to construct the first coordinate sequence;

[0018] Step S4: Generate a first random positive integer i between [1, N], select all coordinates with the first random positive integer as the horizontal coordinate from the first coordinate sequence, and use them to construct a second coordinate sequence of length k+1, and generate a second random integer j between [0, k].

[0019] Step S5: Obtain the position coordinates of the geometric center from the second coordinate sequence (X... i Y m+j The sample units are identified, and the quality risk index P of the sample units is obtained. risk m is a positive integer;

[0020] Step S6: Generate a random number r between [0, 1] and determine P. risk Is it greater than r?

[0021] If not, proceed sequentially from step S4 to step S6;

[0022] If so, the sample unit is added to the submitted sample set, and the number of sample units in the submitted sample set is checked to see if it is equal to n: if so, proceed to step S7; otherwise, proceed to steps S4 to S6 in sequence.

[0023] Step S7: Sort the quality risk index of each sample unit in the submitted sample set from high to low, call the automated testing tool and test the sample units in sequence according to the sorting results, obtain the test conclusion of whether the sample unit is qualified or not, and send it to the judgment module.

[0024] Step S8: Once the decision module receives a non-compliant test result, it outputs an instruction to refuse to accept the current inspection batch. If no non-compliant test result is received after testing every sample unit in the entire sample set, it outputs an instruction to accept the current inspection batch.

[0025] Further, step S2 calculates and obtains the quality risk index for each sample unit in the inspection batch, specifically including:

[0026] Step S21: Obtain the value range A of the sample cell;

[0027] Step S22: Obtain the point feature scale U of the sample unit p Line element scale U l Scale of dough elements U a ;

[0028] Step S23: Obtain the water network complexity E of the sample unit. h and the complexity E of the traffic data network r ;

[0029] Step S24: Obtain the correlation strength H between residential areas and facilities and water network in the sample unit. j And the element correlation strength R between residential areas and facilities and transportation data networks j ;

[0030] Step S25: Set the value range scale A of the sample unit and the point feature scale U. p Line element scale U l Scale of dough elements U a Water network complexity E h Traffic data network complexity E r The intensity of the correlation between residential areas and facilities and water networks (H) j And the element correlation strength R between residential areas and facilities and transportation data networks j Perform normalization;

[0031] Step S26: Normalize A and U p U l U a E h E r H j R j Multiply the results to obtain the quality risk index of the sample unit.

[0032] Further, obtaining the range A of the sample cells in step S21 specifically includes:

[0033] Obtain the set of non-system fields consisting of non-system fields in the sample unit;

[0034] By calculating the value range of each field in the set of non-systematic fields, we obtain a set of value ranges for each field.

[0035] Take the average of the elements in the set of value ranges for the field, and use that average as the value range A of the sample unit.

[0036] Furthermore, in step S22, the point feature scale U of the sample unit is obtained. p Line element scale U lScale of dough elements U a Specifically, it includes:

[0037] The number of point features per unit area, the length of line features per unit area, and the boundary line length of polygon features per unit area in the sample unit are statistically analyzed and calculated, and these are respectively used as the point feature scale U of the sample unit. p Line element scale U l Scale of dough elements U a .

[0038] Furthermore, in step S23, the water network complexity E of the sample unit is obtained. h and the complexity E of the traffic data network r Specifically, it includes:

[0039] The number of nodes N in the water network in the statistical sample unit ph And the number N of the traffic data network in the statistical sample unit. pr ;

[0040] Count the number N edges connected to the nodes in the water network. eh And to count the number N edges connected to nodes in the traffic data network. er ;

[0041] The computational complexity E of the water network h =N eh / N ph And the computational complexity E of the traffic data network. r =N er / N pr .

[0042] Furthermore, in step S24, the correlation strength H between residential areas and facilities and the water network in the sample unit is obtained. j And the element correlation strength R between residential areas and facilities and transportation data networks j Specifically, it includes:

[0043] Acquire data on residential areas and facilities, water networks, and transportation networks within the sample units;

[0044] Calculate the minimum distance from each element in the water network to all elements in the residential areas and facilities, and determine whether the minimum distance is less than the association distance threshold d. If it is less, then perform the first association count C. th The count increases by 1;

[0045] Calculate the minimum distance from each element in the traffic data network to all elements in the residential area and facilities, and determine whether the minimum distance is less than the association distance threshold d. If it is less, then perform the second association count C. tr The count increases by 1;

[0046] Calculate the element correlation strength H between residential areas and facilities and water network. j =First association count C th / Number of elements in the water network;

[0047] Calculate the element correlation strength R between residential areas and facilities and the transportation data network. j =Second association number C tr / Number of elements in the transportation data network;

[0048] The association distance threshold d is preset.

[0049] Furthermore, in step S25, the normalization method used is the minimum-maximum method.

[0050] Furthermore, the automated detection tool is one of ArcGIS, QGIS, or FME.

[0051] Furthermore, the number of samples n submitted for testing is obtained by multiplying N by a preset ratio.

[0052] On the other hand, the present invention also discloses a computer-readable storage medium storing a computer program / instructions that, when executed by a processor or compiled, implement the steps of the method as described in any of the preceding claims.

[0053] The technical solution of this invention has one or more of the following beneficial technical effects:

[0054] (1) It can effectively improve the accuracy of defect detection and the reliability of inspection conclusions of two-dimensional vector geographic databases.

[0055] (2) It not only ensures the high probability of high-quality risk sample units being sampled, but also ensures the randomness and uniformity of the spatial distribution of sample units, avoiding systematic errors and providing a rigorous theoretical basis for the objectivity and accuracy of defect detection.

[0056] (3) Based on the quality risk index in the submitted sample set, the sample units are inspected from high to low. If any unqualified sample unit is detected, the inspection is stopped and the current inspection batch is deemed unqualified. The higher the probability of unqualified high-risk sample units, the lower the expected inspection time of the inspection batch, the lower the overall inspection time, and the higher the inspection efficiency.

[0057] Furthermore, other beneficial effects of the present invention will be mentioned in the specific embodiments. Attached Figure Description

[0058] Figure 1 This is a system architecture diagram of the present invention;

[0059] Figure 2This is a schematic diagram of the defect detection module of the present invention;

[0060] Figure 3 This is a flowchart of a two-dimensional vector geographic database defect detection method;

[0061] Figure 4 This is a flowchart illustrating the method for obtaining the quality risk index for each sample unit. Detailed Implementation

[0062] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0063] To facilitate a clear description of the technical solutions in the embodiments of the present invention, the terms "first" and "second" are used to distinguish identical or similar items with essentially the same function and effect. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order.

[0064] Terminology Explanation:

[0065] Sample Unit: A sample unit is the smallest data unit for implementing defect detection in a two-dimensional vector geographic database. In a two-dimensional vector geographic database, it refers to an independent area defined according to specific spatial division rules (such as a standard map sheet). Each sample unit contains all the topographic features (such as contour lines, landform feature points, etc.) and land feature features (such as roads, water networks, residential areas, etc.) within that area, and can independently carry the complete data information required for defect detection.

[0066] An inspection lot is a collection of several sample units grouped together according to certain logic (such as the same production batch, the same data scale, the same geographical coverage, the same map sheet, etc.), serving as the overall object for defect sampling inspection. The defect detection sampling process involves selecting a certain number of sample units from the inspection lot and using them as the objects for specific defect detection. The overall quality level of the inspection lot is inferred from the inspection results of the selected samples.

[0067] Digital Line Graphic (DLG) is a core vector data product in the field of geographic information. It performs layered vectorization of geographic elements such as terrain and features, while preserving the spatial relationships and attribute information of these elements. It serves as a crucial data carrier connecting topographic maps with geographic information systems and computer-aided design (CAD). DLG eliminates redundant auxiliary information from topographic maps (such as some map border annotations and non-core decorative elements), retaining only core elements with geographic analysis and application value. It accurately describes the spatial location, shape, and interrelationships of elements through vector data structures (points and lines), and adds attribute information (such as road grade, land use type, and water system name). A DLG with a database can be considered a sample unit in this invention (a DLG without a database cannot be considered a sample unit).

[0068] In this invention, the term "geospatial data" has the broadest scope, encompassing all data related to geospatial data, and can be divided into vector data and raster data according to data format; "digital line map" is a subclass of geospatial data, which belongs only to geospatial data in vector format; "two-dimensional vector geodatabase" is a storage medium or container, and digital line map is one of the specific contents within the container.

[0069] Figure 1 This is a system architecture diagram of the present invention. The system architecture includes a computer-readable storage medium, a processor, and input / output (I / O) interfaces and I / O devices. The computer-readable storage medium stores a two-dimensional vector geographic database and a defect detection module constructed in the form of computer software. The computer-readable storage medium, the processor, and the I / O interfaces can establish information connections or / and interact through a system bus.

[0070] The defect detection module can be implemented in the form of a computer program / instruction. When the computer program / instruction is executed by a processor or compiled and executed, it implements the steps of the two-dimensional vector geographic database defect detection method of the present invention.

[0071] Figure 2 This is a schematic diagram of the defect detection module of the present invention. In this invention, the defect detection module includes automated detection tools, such as ArcGIS, QGIS, and FME (data processing software), which can batch detect quantifiable errors such as coordinate deviations, geometric errors, topological conflicts, and missing attributes. By calling these automated detection tools, the efficiency of quality inspection can be improved in this invention. This invention does not limit the specific type of automated detection tool.

[0072] In this invention, the defect detection module includes a sample unit screening module. The sample unit screening module receives inspection batches and generates a set of samples to be submitted to the automated testing tool. If the automated testing tool outputs a non-conforming result, the decision module included in the defect detection module outputs an instruction to reject the current inspection batch. If all samples in the submitted sample set are tested by the automated software and all output a conforming result, the decision module outputs an instruction to accept the current inspection batch. In other words, this invention performs a zero-rejection sampling inspection on non-conforming sample units.

[0073] Through practice, the inventors have realized that geospatial data with different spatial locations and different features of land cover have significant differences in the probability of quality problems and the value of data application. This difference stems from the spatial heterogeneity of geospatial data itself—that is, the distribution density, type combination, and processing difficulty of geographical elements will vary with spatial location, which in turn directly affects the level of data quality risk and practical application value.

[0074] Taking topographic map geospatial data products of the same area and standard scale as examples, a comparative analysis is conducted on such products from the eastern coastal region and the northwestern region. Among them, the topographic map products of the eastern coastal region contain a high-density and complex road network (such as a multi-level road network system formed by urban expressways, main roads and branch roads), water network (such as a water network system composed of rivers, artificial canals, lakes and supporting water conservancy facilities), and dense residential areas (such as multi-story building clusters in urban built-up areas and concentrated residential areas in towns). The above-mentioned geographical elements are intertwined and diverse in spatial distribution, which significantly increases the difficulty of data collection, the complexity of element editing and the cost of quality control of topographic map products in this region. Its application scenarios in urban planning, traffic management, water conservancy scheduling and other fields are more critical, so it has higher application value. At the same time, the requirements for quality indicators such as data accuracy and topological consistency are also more stringent, and it should be the focus of attention in the process of geospatial data quality inspection.

[0075] In contrast, topographic map products in the Northwest region exhibit a sparse road network (mainly national and provincial highways with low coverage of branch roads), rare natural waterways, and scattered artificial water conservancy facilities. Settlements are primarily small-scale villages with dispersed spatial distribution. The uniformity and sparse distribution of geographical elements in this region simplify the data production process for topographic map products, significantly reducing data processing complexity. Consequently, their application in practical applications is less dependent on specific scenarios and less critical, making their application value lower than similar products from the eastern coastal regions. Therefore, the requirements for data quality indicators can be appropriately relaxed, and they should not be prioritized for quality inspection.

[0076] In addition to the differences between the eastern coastal areas and the northwestern regions mentioned above, in the practice of quality inspection of geospatial data, there are also such quality risks and application value differences caused by different spatial positions and feature distribution characteristics in the geospatial data products of urban and rural areas, plain and mountainous areas: For example, the geospatial data of urban areas contains elements such as dense buildings, underground pipelines, and transportation facilities, with high data complexity and great application value, and thus requires key inspection; rural areas mainly consist of agricultural land and scattered residential areas, with lower data complexity and application value; plain areas have flat terrain and regular distribution of geographical elements, with low data production difficulty and low quality risk; mountainous areas have complex terrain and landforms (such as dense contour lines and large slope changes), and may contain special features (such as mountain vegetation and geological disaster points), with great difficulty in data collection and processing and high quality risk, which should be the key focus of inspection.

[0077] Figure 3 It is a flowchart of the defect detection method for a two-dimensional vector geodatabase. Based on this discovery, the sample unit screening module in the present invention is configured to perform the following steps, ensuring a high probability of sampling for high-quality risk products and also ensuring the randomness and uniformity of the distribution of sample units in the sample space:

[0078] Step S1: The sample unit screening module receives the inspection batch, obtains the number N of sample units in the inspection batch, and configures the number n of samples to be sent for inspection.

[0079] Both N and n here are positive integers, and n < N. The number n of samples to be sent for inspection here can be a preset value or calculated according to a preset ratio × N. The present invention is not limited to this and does not take this as a limitation.

[0080] A large number of sample units included in the inspection batch in the present invention can be surveying and mapping geographic information products from "real geographical entities" to "digital vector expressions" collected / obtained through various channels. The present invention is not limited to surveying and mapping geographic information products obtained in a certain specific form.

[0081] Step S2: Calculate and obtain the quality risk index of each sample unit in the inspection batch.

[0082] The risk quality index proposed by the present invention is an indicator indicating the quality risk of sample units.

[0083] Figure 4 It is a flowchart of the method for obtaining the quality risk index of each sample unit. In order to objectively, accurately, and comprehensively measure the quality risk of each sample unit, the present invention proposes the following specific method for obtaining the quality risk index:

[0084] Step S21: Obtain the value range A of the sample unit. A sample unit includes various geographic information elements, such as roads in a transportation data network. Roads have attributes / fields such as code, length, and width. For example, the width of a road may have multiple values.

[0085] In this invention, the value range of a field refers to the total number of values ​​that the field can take within its sample unit. A larger value range indicates richer attribute information within the sample unit, resulting in higher production difficulty and higher quality risk under the same conditions.

[0086] To obtain the value range of a sample cell, first obtain the set of non-system fields [F1, F2, ..., F] that are non-system fields in the sample cell. u Non-system fields refer to user-defined fields in the data generation software, while system fields are fields generated by default by the software. The system field set includes feature unique identifier fields, geometry type fields, system length fields, system area fields, and other non-user-defined fields.

[0087] Next, the non-system field set [F1, F2, ..., F] is statistically analyzed. u The value range of each field in the dataset is used to obtain the set of value ranges for the fields [V1, V2, ..., V]. u Then, the range size A of the sample cells is calculated to be equal to the average of the elements in the set of range sizes of the field, where u is a positive integer.

[0088] Step S22: Obtain the point feature scale U of the sample unit p Line element scale U l Scale of dough elements U a The more point features, line features, and boundary lines of surface features within a unit area of ​​a sample cell, the more numerous and denser the geometric elements within that sample cell. Under the same conditions, this results in greater production difficulty and a higher probability of quality risks.

[0089] Specifically, the spatial area S of the sample unit can be obtained. Then, based on the spatial area S, the number of point features per unit area, the length of line features per unit area, and the boundary line length of surface features per unit area can be statistically calculated to express the point feature scale U of the sample unit. p Line element scale U l Scale of dough elements U a .

[0090] Step S23: Obtain the water network complexity E of the sample unit. h and the complexity E of the traffic data network r .

[0091] Water networks and transportation data networks require handling connectivity and network topology issues during production, and connectivity and network topology are also important targets for defect detection. This invention, in addition to measuring the quality risk index from the perspective of element scale, also considers network complexity.

[0092] In a network, the more nodes and edges connecting them, the more complex the network, the more difficult it is to process, and the higher the risk of quality defects. This invention specifically obtains the water network complexity E through the following method. h and the complexity E of the traffic data network r .

[0093] First, count the number of nodes in the water network within the sample unit, denoted as N. ph ; and the number of nodes in the traffic data network within the statistical sample unit, denoted as N. pr Next, count the number of edges connecting the nodes in the water network, denoted as N. eh ; and the number of edges connecting nodes in the statistical traffic data network, denoted as N. er .

[0094] Finally, the complexity E of the water network is calculated. h =N eh / N ph And the computational complexity E of the traffic data network. r =N er / N pr In this invention, the water network complexity E h and the complexity E of the traffic data network r , is the average number of connected edges in a network, and is a way to express network complexity.

[0095] Step S24: Obtain the correlation strength H between residential areas and facilities and water network in the sample unit. j And the element correlation strength R between residential areas and facilities and transportation data networks j .

[0096] Residential areas and facilities, water networks, and transportation data networks are all geographic information elements. Residential areas and facilities along with water networks, and residential areas and facilities along with transportation data networks, are crucial components in the production process of two-dimensional vector geographic databases, and also areas prone to quality issues. The more relationships that need to be processed, the greater the generation difficulty and the higher the quality risk.

[0097] To determine whether relationships between different elements need to be processed, this invention employs a threshold judgment method. That is, if the distance between elements is less than a threshold, then the relationship between the elements needs to be processed; otherwise, the relationship between the elements does not need to be processed.

[0098] Specifically, a correlation distance threshold d can be preset. Next, the residential areas and facilities, water network, and transportation network data from the sample units are obtained. The minimum distance from each element in the water network to all elements in the residential areas and facilities is calculated. It is then determined whether the minimum distance is less than the correlation distance threshold d. If it is less, the first correlation iteration C is performed. th The count is increased by 1; and the minimum distance from each element in the traffic data network to all elements in the residential area and facilities is calculated, and it is determined whether the minimum distance is less than the association distance threshold d. If it is less, the second association count C is performed. tr The count increases by 1.

[0099] Then, using the first association number C mentioned above... th Dividing by the number of elements in the water network yields the element correlation strength H between settlements and facilities and the water network. j Then, using the second association number C mentioned above tr Dividing by the number of elements in the transportation data network yields the element correlation strength R between residential areas and facilities and the transportation data network. j In this invention, H j and R j Each represents the average number of associations, expressing the strength of the association between different sets of elements.

[0100] Step S25: Set the value range scale A of the sample unit and the point feature scale U. p Line element scale U l Scale of dough elements U a Water network complexity E h Traffic data network complexity E r The intensity of the correlation between residential areas and facilities and water networks (H) j And the element correlation strength R between residential areas and facilities and transportation data networks j Normalize.

[0101] The specific normalization method can follow the well-known Min-Max method, which will not be elaborated here. After normalization, the dimensional differences between different types of data can be eliminated.

[0102] Step S26: Normalize A and U p U l U a E h E r H j R j Multiply the results to obtain the quality risk index of the sample unit.

[0103] Through the aforementioned steps S21 to S26, the solution of the present invention can obtain an objective, accurate, and comprehensive quantification of the risks existing in the sample units, providing a favorable foundation for the present invention.

[0104] Step S3: Obtain the position coordinates of the geometric center of each sample unit to construct the first coordinate sequence.

[0105] For example, the first coordinate sequence is [(X1, Y1), (X2, Y2), ..., (X... N Y N ], it is an ordered set.

[0106] Step S4: Generate a first random positive integer i between [1, N], select all coordinates whose horizontal coordinate index is the first random positive integer from the first coordinate sequence, and use them to construct a second coordinate sequence of length k+1, and generate a second random integer j between [0, k].

[0107] We can assume the first random positive integer is i, then (X) i Y i Let be the position coordinates of the geometric center of the i-th sample unit, where i is a positive integer between [1, n]. Then, the second coordinate sequence is [(X... i Y m (X) i Y m+1 ), ..., (X i Y m+k ]], where m≥1, m+k≤N, and m is a positive integer. Assume the second random integer is j, and j is an integer between [0, k] (j is allowed to be equal to 0).

[0108] Step S5: Obtain the position coordinates of the geometric center as (X... i Y m+j The sample unit (denoted as P) is selected, and the quality risk index (denoted as P) of the sample unit P is obtained. risk ).

[0109] Step S6: Generate a random number r between [0, 1] and determine P. risk Is it greater than r? If not, proceed to steps S4 to S6 in sequence. If yes, add the sample unit to the submitted sample set and check if the number of sample units in the submitted sample set is equal to n. If yes, proceed to step S7; otherwise, proceed to steps S4 to S6 in sequence.

[0110] Here, the random number r is a decimal used for comparison with P. risk The size determines whether the current sample unit should be included in the sample.

[0111] Step S7: Sort the quality risk index of each sample unit in the submitted sample set from high to low, call the automated testing tool and test the sample units in sequence according to the sorting results, obtain the test conclusion of whether the sample unit is qualified or not, and send it to the decision module.

[0112] Since each test only examines the sample unit with the highest quality risk index among the remaining sample units, this step can minimize the overall testing time and shorten the valuable testing cycle.

[0113] Step S8: Once the decision module receives a non-compliant test result, it outputs an instruction to refuse to accept the current inspection batch. If no non-compliant test result is received after testing every sample unit in the entire sample set, it outputs an instruction to accept the current inspection batch.

[0114] Finally, in several examples, the inventors used the defect detection system developed in this invention to verify the technical effect of the invention with actual data, and the sample unit was a DLG with a pre-built library.

[0115] In the first example, the number of sample units N in the inspection batch is 1000, and the number of samples submitted for testing n is 25 (i.e., 25 sample units must be sampled for testing). Table 1 is the final set of samples submitted for testing in this example. The average quality risk of the sample set submitted for testing in this invention is calculated to be 0.7816.

[0116] Table 1: Final set of submitted samples in one embodiment of the present invention

[0117]

[0118] To verify and compare the differences between the present invention and the prior art, Table 2 presents the sample set obtained by conventional methods. The average quality risk of the sample set in this comparative example was calculated to be 0.6855. This indicates that in this example, the present invention can extract sample units with a 14.02% higher probability than the prior art.

[0119] Table 2: A set of submitted samples obtained based on existing technology

[0120]

[0121] In the second example, the number of sample units N in the inspection batch is 1000, and the number of submitted samples n is 50. Table 3 is the final set of submitted samples in this example. The average quality risk of the submitted sample set of the present invention is calculated to be 0.8546.

[0122] Table 3: Final sample set submitted for testing in another embodiment of the present invention

[0123]

[0124] To verify and compare the differences between the present invention and the prior art, Table 4 presents another set of submitted samples obtained using conventional methods. The average quality risk of this set of submitted samples in the comparative example was calculated to be 0.6926. This indicates that in this example, the present invention can extract sample units with a 23.39% higher probability than the prior art.

[0125] Table 4: Another set of submitted samples obtained based on existing technology

[0126]

[0127] In the third example, the number of sample units N in the inspection batch is 1000, and the number of submitted samples n is 75. Table 5 shows the final set of submitted samples in this example. The average quality risk of the submitted sample set of the present invention is calculated to be 0.8428.

[0128] Table 5: Final Set of Submitted Samples in Yet Another Example of the Invention

[0129]

[0130] To verify and compare the differences between the present invention and the prior art, Table 6 presents another set of submitted samples obtained using conventional methods. The average quality risk of this set of submitted samples in the comparative example was calculated to be 0.7497. This indicates that in this example, the present invention can extract sample units with a 12.42% higher probability than the prior art.

[0131] Table 6: Another set of submitted samples obtained based on existing technology

[0132]

[0133] In summary, the above examples also verify the effectiveness of the technical solution of the present invention. The present invention can effectively improve the accuracy of defect detection and the reliability of inspection conclusions of two-dimensional vector geographic databases.

[0134] To better illustrate the present invention, numerous specific details have been provided in the detailed embodiments described above. Those skilled in the art should understand that the present invention can be practiced even without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art have not been described in detail in order to highlight the main points of the present invention.

[0135] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A method for detecting defects in a two-dimensional vector geographic database, characterized in that, include: Step S1: The sample unit screening module receives the test batch, obtains the number N of sample units in the test batch, and configures the number of samples to be sent for testing, n. Step S2: Calculate and obtain the quality risk index of each sample unit in the inspection batch; Step S3: Obtain the position coordinates of the geometric center of each sample unit to construct the first coordinate sequence; Step S4: Generate a first random positive integer i between [1, N], select all coordinates with the first random positive integer as the horizontal coordinate from the first coordinate sequence, and use them to construct a second coordinate sequence of length k+1, and generate a second random integer j between [0, k]. Step S5: Obtain the position coordinates of the geometric center from the second coordinate sequence (X... i Y m+j The sample units are identified, and the quality risk index P of the sample units is obtained. risk m is a positive integer; Step S6: Generate a random number r between [0, 1] and determine P. risk Is it greater than r? If not, proceed sequentially from step S4 to step S6; If so, the sample unit is added to the submitted sample set, and the number of sample units in the submitted sample set is checked to see if it is equal to n: if so, proceed to step S7; otherwise, proceed to steps S4 to S6 in sequence. Step S7: Sort the quality risk index of each sample unit in the submitted sample set from high to low, call the automated testing tool and test the sample units in sequence according to the sorting results, obtain the test conclusion of whether the sample unit is qualified or not, and send it to the judgment module. Step S8: Once the decision module receives a non-compliant test result, it outputs an instruction to refuse to accept the current inspection batch. If no non-compliant test result is received after testing every sample unit in the entire sample set, it outputs an instruction to accept the current inspection batch. Specifically, step S2 involves calculating and obtaining the quality risk index for each sample unit in the inspection batch, including: Step S21: Obtain the value range A of the sample cell; Step S22: Obtain the point feature scale U of the sample unit p Line element scale U l Scale of dough elements U a ; Step S23: Obtain the water network complexity E of the sample unit. h and the complexity E of the traffic data network r ; Step S24: Obtain the correlation strength H between residential areas and facilities and water network in the sample unit. j And the element correlation strength R between residential areas and facilities and transportation data networks j ; Step S25: Set the value range scale A of the sample unit and the point feature scale U. p Line element scale U l Scale of dough elements U a Water network complexity E h Traffic data network complexity E r The intensity of the correlation between residential areas and facilities and water networks (H) j And the element correlation strength R between residential areas and facilities and transportation data networks j Perform normalization; Step S26: Normalize A and U p U l U a E h E r H j R j Multiply the results to obtain the quality risk index of the sample unit.

2. The two-dimensional vector geographic database defect detection method according to claim 1, characterized in that: Step S21 involves obtaining the range A of the sample cells, specifically including: Obtain the set of non-system fields consisting of non-system fields in the sample unit; By calculating the value range of each field in the set of non-systematic fields, we obtain a set of value ranges for each field. Take the average of the elements in the set of value ranges for the field, and use that average as the value range A of the sample unit.

3. The two-dimensional vector geographic database defect detection method according to claim 2, characterized in that: In step S22, the point feature scale U of the sample unit is obtained. p Line element scale U l Scale of dough elements U a Specifically, it includes: The number of point features per unit area, the length of line features per unit area, and the boundary line length of polygon features per unit area in the sample unit are statistically analyzed and calculated, and these are respectively used as the point feature scale U of the sample unit. p Line element scale U l Scale of dough elements U a .

4. The two-dimensional vector geographic database defect detection method according to claim 3, characterized in that: In step S23, the water network complexity E of the sample unit is obtained. h and the complexity E of the traffic data network r Specifically, it includes: The number of nodes N in the water network in the statistical sample unit ph And the number N of the traffic data network in the statistical sample unit. pr ; Count the number N edges connected to the nodes in the water network. eh And to count the number N edges connected to nodes in the traffic data network. er ; The computational complexity E of the water network h =N eh / N ph And the computational complexity E of the traffic data network. r =N er / N pr .

5. The two-dimensional vector geographic database defect detection method according to claim 4, characterized in that: In step S24, the correlation strength H between residential areas and facilities and water networks in the sample unit is obtained. j And the element correlation strength R between residential areas and facilities and transportation data networks j Specifically, it includes: Acquire data on residential areas and facilities, water networks, and transportation networks within the sample units; Calculate the minimum distance from each element in the water network to all elements in the residential areas and facilities, and determine whether the minimum distance is less than the association distance threshold d. If it is less, then perform the first association count C. th The count increases by 1; Calculate the minimum distance from each element in the traffic data network to all elements in the residential area and facilities, and determine whether the minimum distance is less than the association distance threshold d. If it is less, then perform the second association count C. tr The count increases by 1; Calculate the element correlation strength H between residential areas and facilities and water network. j =First association count C th / Number of elements in the water network; Calculate the element correlation strength R between residential areas and facilities and the transportation data network. j =Second association number C tr / Number of elements in the transportation data network; The association distance threshold d is preset.

6. The two-dimensional vector geographic database defect detection method according to claim 5, characterized in that: In step S25, the normalization method used is the minimum-maximum method.

7. The two-dimensional vector geographic database defect detection method according to claim 6, characterized in that: The automated detection tool is one of ArcGIS, QGIS, or FME.

8. The two-dimensional vector geographic database defect detection method according to claim 7, characterized in that: The number of samples n submitted for testing is obtained by multiplying N by a preset ratio.

9. A computer-readable storage medium storing a computer program / instructions, characterized in that, When the computer program / instructions are executed by a processor or compiled and executed, they implement the steps of the method according to any one of claims 1-8.

Citation Information

Patent Citations

  • Digital line drawing quality inspection sampling method using spatial heterogeneity

    CN118071208A

  • A surveying and mapping geographic information data management method and system based on cloud computing

    CN119557385B

  • Software defect and complexity incidence relation analysis method based on machine learning

    CN111338972A

  • Efficient geospatial search coverage tracking for detection of dangerous, valuable, and / or other objects dispersed in a geospatial area

    US20240015690A1