Automatic quality inspection system and method for multi-source place name data
Through the automated quality inspection system for multi-source place name data, and by utilizing geographic proximity and deep learning semantic analysis technology, the problems of homonyms and semantic ambiguity in the quality inspection of multi-source place name data are solved, thereby improving data quality and integration efficiency and supporting urban construction planning.
Patent Information
- Application Number
- CN202510368648.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-27
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-03-27
AI Technical Summary
The quality inspection methods for multi-source place name data in existing technologies are inefficient and difficult to guarantee accuracy. They are unable to effectively handle problems such as the same name in different places, the same place with different names, and semantic ambiguity, resulting in low data quality and integration efficiency.
An automated quality inspection system for multi-source place name data is used. After preliminary screening based on geographic proximity, semantic parsing and implicit semantic query interactive analysis are performed in combination with deep learning semantic analysis technology to determine the entity semantic alignment between place name data.
Effectively solve the problems of homonymous places, different names for the same place, and semantic ambiguity, improve the quality and integration efficiency of place name data, and provide accurate place name data support for applications such as urban construction planning.
Smart Images

Figure CN119903196B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data analysis technology, and more specifically, to an automated quality inspection system and method for multi-source place name data. Background Art
[0002] With the rapid development of information technology, the demand for place name data in numerous fields, including geographic information systems (GIS), intelligent navigation, and location-based services, is growing and demanding. As a crucial geographic information resource, place name data is widely distributed across multiple data sources, including government departments (such as civil affairs and surveying and mapping), businesses (such as map service providers), and various open source data platforms.
[0003] However, place name data from different data sources vary significantly in terms of content detail and update frequency. For example, government place name data may focus on basic information such as administrative divisions and standardized place names, while corporate place name data may focus more on commercially valuable information such as commercial facilities and tourist attractions, and may also contain varying degrees of information redundancy or omission. Regarding update frequency, open source data platforms, due to their reliance on user contributions, may have inconsistent data update schedules, making it difficult to guarantee data timeliness.
[0004] Due to the heterogeneity of data sources, problems such as the same name in different places, different names for the same place, and semantic ambiguity frequently occur, which seriously restricts the effective integration and application of place name data. Existing quality inspection methods mainly rely on manual review or simple rule matching. Manual review is extremely inefficient. Faced with massive multi-source place name data, it requires a lot of manpower, material resources and time, and is easily affected by human factors, making it difficult to ensure the accuracy and consistency of quality inspection results. Simple rule matching methods, such as those based on full matching or partial string matching of place name text, cannot effectively handle semantic changes such as synonyms, abbreviations, and aliases of place names, and it is also difficult to explore potential semantic associations between place names. For example, "Beijing" and "Beijing" have different text forms, but semantically point to the same entity. Traditional rule matching methods often cannot recognize this relationship, resulting in data matching errors, affecting data quality and subsequent application effects.
[0005] Therefore, an optimized automatic quality inspection system and method for multi-source place name data is expected. Summary of the Invention
[0006] In order to solve the above technical problems, the present application is proposed. The embodiment of the present application provides a multi-source place name data automated quality inspection system and method, which first accesses place name data from multiple data sources, and performs preliminary matching screening on the accessed place name data based on geographical proximity, and then further introduces semantic analysis technology based on deep learning, and performs semantic parsing and implicit semantic query interaction analysis on the two place name data in the candidate pair to mine the potential semantic association between the two, determine the entity semantic alignment between the place name data, and then determine whether the two belong to the same place name entity based on the comparison of the entity semantic alignment with the preset threshold. In this way, problems such as the same name in different places, different names for the same place, and semantic ambiguity can be effectively solved, the quality and integration efficiency of place name data can be improved, and accurate place name data support can be provided for applications such as urban construction planning.
[0007] According to one aspect of the present application, a method for automated quality inspection of multi-source place name data is provided, comprising:
[0008] Accessing place name data from multiple data sources to obtain a set of place name data, wherein the place name data includes place name, coordinates and attribute information;
[0009] Based on geographical location proximity, screening the set of place name data for preliminary candidate pairs to obtain a set of preliminary candidate pairs of place name data;
[0010] Extracting preliminary candidate pairs of place name data to be quality checked from the set of preliminary candidate pairs of place name data, wherein the preliminary candidate pairs of place name data to be quality checked include first place name data and second place name data;
[0011] Calculating an entity semantic alignment between the first place name data and the second place name data, wherein calculating the entity semantic alignment between the first place name data and the second place name data includes: performing a latent semantic query space interactive alignment analysis on the first place name data and the second place name data to obtain the entity semantic alignment;
[0012] Based on the comparison between the entity semantic alignment and a preset threshold, it is determined whether the first place-name data and the second place-name data belong to the same place-name entity.
[0013] Preferably, performing latent semantic query space interactive alignment analysis on the first place name data and the second place name data to obtain the entity semantic alignment includes:
[0014] Performing structured mapping on the first place name data and the second place name data to obtain a first place name data structured mapping encoding vector and a second place name data structured mapping encoding vector;
[0015] Performing semantic alignment analysis on the first place name data structured mapping encoding vector and the second place name data structured mapping encoding vector based on implicit query space sparse constraints to obtain a first-second place name data semantic interaction encoding feature vector;
[0016] The entity semantic alignment is determined based on the semantic interaction encoding feature vector of the first and second place name data.
[0017] Preferably, performing structured mapping on the first place name data and the second place name data to obtain a first place name data structured mapping encoding vector and a second place name data structured mapping encoding vector further includes:
[0018] The first place name data and the second place name data are semantically encoded using a semantic encoder based on the Bert model to obtain a structured mapping encoding vector of the first place name data and a structured mapping encoding vector of the second place name data.
[0019] Preferably, performing semantic alignment analysis based on implicit query space sparse constraints on the first place name data structured mapping encoding vector and the second place name data structured mapping encoding vector to obtain the first-second place name data semantic interaction encoding feature vector includes:
[0020] Performing implicit query interaction based on a local semantic space on the first place name data structured mapping code vector and the second place name data structured mapping code vector to obtain a set of first-second place name data semantic interaction implicit query space joint coding matrices;
[0021] Adaptive aggregation based on implicit query space sparse constraints is performed on the set of the first-second place name data semantic interaction implicit query space joint encoding matrices to obtain the first-second place name data semantic interaction encoding feature vector.
[0022] Preferably, performing implicit query interaction based on a local semantic space on the first place name data structured mapping code vector and the second place name data structured mapping code vector to obtain a set of first-second place name data semantic interaction implicit query space joint coding matrices includes:
[0023] Performing feature phase space reconstruction based on one-dimensional convolution coding on the first place name data structured mapping code vector and the second place name data structured mapping code vector to obtain a set of first place name data local semantic feature code vectors and a set of second place name data local semantic feature code vectors;
[0024] Calculate the implicit query space joint coding matrix between any set of the first place name data local semantic feature coding vectors and the second place name data local semantic feature coding vectors in the set of the first place name data local semantic feature coding vectors and the set of the second place name data local semantic feature coding vectors to obtain the set of the first-second place name data semantic interaction implicit query space joint coding matrices.
[0025] Preferably, performing adaptive aggregation based on the sparse constraint of the implicit query space on the set of the joint encoding matrices of the semantic interaction between the first and second place name data to obtain the semantic interaction encoding feature vector of the first and second place name data includes:
[0026] performing redundant interaction topological reduction on each first-second place name data semantic interaction implicit query space joint encoding matrix in the set of the first-second place name data semantic interaction implicit query space joint encoding matrices to obtain a set of optimized first-second place name data semantic interaction implicit query space joint encoding matrices;
[0027] Calculating the spatial sparse constraint factor of each optimized first-second place name data semantic interaction implicit query space joint encoding matrix in the set of optimized first-second place name data semantic interaction implicit query space joint encoding matrices to obtain a set of first-second place name data semantic interaction spatial sparse constraint factors;
[0028] Based on the set of sparse constraint factors of the first-second place name data semantic interaction space, the set of optimized first-second place name data semantic interaction implicit query space joint coding matrices is adaptively aggregated and encoded to obtain the first-second place name data semantic interaction coding feature vector.
[0029] Preferably, redundant interaction topology reduction is performed on each first-second place name data semantic interaction implicit query space joint encoding matrix in the set of the first-second place name data semantic interaction implicit query space joint encoding matrices to obtain a set of optimized first-second place name data semantic interaction implicit query space joint encoding matrices, including:
[0030] Constructing a first place name data-second place name data semantic interaction phase matrix between the first place name data local semantic feature encoding vector and the second place name data local semantic feature encoding vector;
[0031] Based on the semantic interaction phase matrix of the first place name data and the second place name data, the implicit query space joint coding matrix between the local semantic feature coding vector of the first place name data and the local semantic feature coding vector of the second place name data is optimized and mapped to obtain the optimized first-second place name data semantic interaction implicit query space joint coding matrix.
[0032] Preferably, the sparse constraint factor of the first-second place name data semantic interaction space is the square of the F norm of the optimized first-second place name data semantic interaction implicit query space joint encoding matrix.
[0033] Preferably, determining the entity semantic alignment based on the semantic interaction encoding feature vector of the first and second place name data includes:
[0034] The semantic interaction encoding feature vector of the first and second place name data is input into a decoder-based semantic alignment quality inspection module to obtain the entity semantic alignment degree.
[0035] According to another aspect of the present application, there is provided a multi-source place name data automated quality inspection system, comprising:
[0036] A data access and integration module, configured to access place name data from multiple data sources to obtain a collection of place name data, wherein the place name data includes place name, coordinates, and attribute information;
[0037] a data candidate pair screening module, configured to screen the set of place name data for preliminary candidate pairs based on geographical location proximity to obtain a set of preliminary candidate pairs of place name data;
[0038] a candidate pair extraction module for extracting preliminary candidate pairs of place name data to be inspected from the set of preliminary candidate pairs of place name data, wherein the preliminary candidate pairs of place name data to be inspected include first place name data and second place name data;
[0039] an entity semantic alignment analysis module, configured to calculate an entity semantic alignment degree between the first place name data and the second place name data, wherein calculating the entity semantic alignment degree between the first place name data and the second place name data comprises: performing a latent semantic query space interactive alignment analysis on the first place name data and the second place name data to obtain the entity semantic alignment degree;
[0040] The homonymous entity determination module is used to determine whether the first place name data and the second place name data belong to the same place name entity based on a comparison between the entity semantic alignment and a preset threshold.
[0041] This application has at least the following technical effects:
[0042] Compared with the existing technology, the multi-source place name data automated quality inspection system and method provided by this application first accesses place name data from multiple data sources, and performs preliminary matching screening on the accessed place name data based on geographical proximity. It then further introduces semantic analysis technology based on deep learning, and performs semantic parsing and implicit semantic query interaction analysis on the two place name data in the candidate pair to explore the potential semantic association between the two, determine the entity semantic alignment between the place name data, and then determine whether the two belong to the same place name entity based on the comparison of the entity semantic alignment with a preset threshold. This application can effectively solve problems such as the same name in different places, different names for the same place, and semantic ambiguity, improve the quality and integration efficiency of place name data, and provide accurate place name data support for applications such as urban construction planning. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0044] Figure 1 Flowchart of an automated quality inspection method for multi-source place name data according to an embodiment of the present application.
[0045] Figure 2 Schematic diagram of data flow of the method for automatic quality inspection of multi-source place name data according to an embodiment of the present application.
[0046] Figure 3 This is a flowchart of sub-step S4 of the method for automated quality inspection of multi-source place name data according to an embodiment of the present application.
[0047] Figure 4 This is a flowchart of sub-step S42 of the method for automated quality inspection of multi-source place name data according to an embodiment of the present application.
[0048] Figure 5 This is a flowchart of sub-step S421 of the method for automated quality inspection of multi-source place name data according to an embodiment of the present application.
[0049] Figure 6 Flowchart of sub-step S422 of the method for automated quality inspection of multi-source place name data according to an embodiment of the present application.
[0050] Figure 7 4 is a block diagram of an automated quality inspection system for multi-source place name data according to an embodiment of the present application. DETAILED DESCRIPTION
[0051] As used in this application and the claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not intended to refer to the singular but may include the plural. Generally speaking, the terms "comprises" and "include" only indicate the inclusion of the steps and elements specifically identified, and these steps and elements do not constitute an exclusive list. A method or apparatus may also include other steps or elements.
[0052] Although the present application makes various references to certain modules in the system according to embodiments of the present application, any number of different modules can be used and run on the user terminal and / or server. The modules are illustrative only, and different aspects of the system and method can use different modules.
[0053] Flowcharts are used in this application to illustrate the operations performed by the systems according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, the various steps may be processed in reverse order or simultaneously, as needed. Furthermore, other operations may be added to these processes, or one or more operations may be removed from these processes.
[0054] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.
[0055] It should be noted that all data acquisition actions in this application are carried out in compliance with the relevant data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.
[0056] In response to the technical problems described in the above background technology, this application proposes an automated quality inspection method for multi-source place name data, which first accesses place name data from multiple data sources, and performs preliminary matching and screening on the accessed place name data based on geographical proximity. Then, it further introduces semantic analysis technology based on deep learning, and performs semantic parsing and implicit semantic query interaction analysis on the two place name data in the candidate pair to explore the potential semantic association between the two, determine the entity semantic alignment between the place name data, and then determine whether the two belong to the same place name entity based on the comparison of the entity semantic alignment with the preset threshold. This can effectively solve problems such as the same name in different places, different names for the same place, and semantic ambiguity, improve the quality and integration efficiency of place name data, and provide accurate place name data support for applications such as urban construction planning.
[0057] Figure 1 Flowchart of an automated quality inspection method for multi-source place name data according to an embodiment of the present application. Figure 2Schematic diagram of data flow of the method for automatic quality inspection of multi-source place name data according to an embodiment of the present application. Figure 1 and Figure 2 As shown, the method for automated quality inspection of multi-source place name data includes the following steps: S1, accessing place name data from multiple data sources to obtain a set of place name data, wherein the place name data includes place name names, coordinates and attribute information; S2, based on geographical location proximity, performing preliminary candidate pair screening on the set of place name data to obtain a set of preliminary candidate pairs of place name data; S3, extracting preliminary candidate pairs of place name data to be quality inspected from the set of preliminary candidate pairs of place name data, wherein the preliminary candidate pairs of place name data to be quality inspected include first place name data and second place name data; S4, calculating the entity semantic alignment between the first place name data and the second place name data, wherein calculating the entity semantic alignment between the first place name data and the second place name data includes: performing implicit semantic query space interactive alignment analysis on the first place name data and the second place name data to obtain the entity semantic alignment; S5, determining whether the first place name data and the second place name data belong to the same place name entity based on a comparison between the entity semantic alignment and a preset threshold.
[0058] In the aforementioned automated quality inspection method for multi-source place name data, step S1 accesses place name data from multiple data sources to obtain a collection of place name data. This place name data includes place name names, coordinates, and attribute information. It should be understood that in real-world scenarios, place name data is widely distributed across different departments and platforms, such as civil affairs departments, mapping agencies, and commercial map service providers. This results in significant differences in format, coordinate system, and attribute fields among place name data from different sources. For example, one data source may use the WGS84 coordinate system, while another uses GCJ02. Attribute information may include different fields such as "administrative division code" and "alias." Given the heterogeneity of data sources, problems such as the same name in different places and coordinate offsets are directly caused. Therefore, during the data access phase, an adapter module must be designed to parse the interface protocols of different data sources (such as APIs, databases, and files), extract place name names, coordinates, and attribute information, and achieve standardization through coordinate system conversion (e.g., using the Proj4 library) and field mapping (e.g., standardizing "province code" to "administrative division code"). For example, the "coordinate" field of a data source is stored as a "longitude, latitude" string and needs to be split into floating-point values; the "alias" field of another data source may contain multiple names separated by semicolons and needs to be converted into a list structure, so that the collected multi-source place name data has a unified field definition, coordinate system and storage format, which reduces the complexity of subsequent steps and improves data consistency.
[0059] During the specific implementation process, considering that each data source may adopt different interface protocols, such as API, database or file format, the adapter needs to have strong parsing capabilities to extract the required key information. For example, when processing place name data, if the data is provided in the form of an API, it is necessary to call the corresponding API interface and configure the parameters according to its documentation to ensure that the required place name, coordinates and attribute information can be correctly obtained. For those place name data stored in the form of databases, such as special systems used within certain enterprises, the adapter module must realize the connection with these databases and extract relevant data by writing SQL query statements. When encountering data sources in the form of files, whether it is CSV, JSON or XML format, it is necessary to develop corresponding parsing logic to convert unstructured or semi-structured data into an actionable information set.
[0060] In addition, due to the differences in coordinate systems among various data sources, for example, one data source may use the WGS84 coordinate system while another uses GCJ02, which requires coordinate system conversion during the access process of place name data. The Proj4 library or other similar tools can be used to efficiently complete the conversion from one coordinate system to another. This conversion not only ensures the consistency of coordinate data, but also lays the foundation for subsequent preliminary matching and screening based on geographic proximity. At the same time, the standardization of attribute information cannot be ignored. Data from different sources may have differences in field naming and content such as "administrative division code" and "alias". Therefore, it is necessary to formulate unified field mapping rules, such as unifying "province code" into "administrative division code", so that the collected multi-source place name data can achieve consistency in field definitions.
[0061] To further improve data consistency and usability, attention should also be paid to the standardization of place names. For example, the "coordinates" field in some place name data exists as a string of "longitude, latitude". In this case, it needs to be split into floating-point values to facilitate calculation and analysis. In cases where multiple aliases are included, such as a data source that lists several names separated by semicolons in the "Alias" field, it should be converted into a list structure for subsequent processing. This processing method not only simplifies the data structure but also improves the efficiency of data retrieval and comparison.
[0062] Ensuring data quality is crucial throughout the entire data access process. To this end, in addition to the technical measures mentioned above, automated testing scripts can be introduced to verify each batch of data received, checking for omissions or errors. For example, pre-set rules can be used to check whether coordinate values fall outside a reasonable range or confirm whether place names conform to specific formatting requirements. If any issues are discovered, relevant personnel are immediately notified and corrected, ensuring that the resulting place name data set is both accurate and reliable.
[0063] In the above-mentioned automated quality inspection method for multi-source place name data, step S2, based on geographic proximity, screens the set of place name data for preliminary candidate pairs to obtain a set of preliminary candidate pairs of place name data. It should be understood that different names for the same entity (such as "Beijing" and "Beijing") must have similar geographic locations. Therefore, based on geographic proximity, this application screens out geographically similar place name data pairs as preliminary candidate pairs of place name data by calculating the geographic distance between the place name data, thereby significantly reducing the processing load of subsequent data deep alignment analysis. In specific implementations, distance metrics such as Euclidean distance, Manhattan distance, or Huffman distance can be used, and a reasonable distance threshold can be set to traverse and compare the set of place name data, screening out two place name data with a relative distance less than the threshold to construct preliminary candidate pairs of place name data. For example, first, a spatial index is established for the place name coordinates, and GeoHash is used to encode the longitude and latitude into strings. Strings with prefix matching represent geographically adjacent areas. For each place name, we search for other place names with the same GeoHash prefix, calculate the Euclidean or Haversine distance, and retain pairs with a distance less than a preset threshold (e.g., 1 kilometer). For example, the straight-line distance between "Beijing Municipal Government (116.407°E, 39.904°N)" and "Beijing Building (116.461°E, 39.926°N)" is approximately 5 kilometers. If the threshold is set to 10 kilometers, these two places will be retained as candidate pairs, thus resolving the issue of having the same name but different locations.
[0064] During implementation, an efficient spatial index structure must be constructed to support rapid retrieval and comparison. By converting the geographic coordinates of all place name data into a unified format and ensuring coordinate system consistency, this lays the foundation for subsequent distance calculations. Spatial indexing techniques such as R-trees or quadtrees can effectively reduce the search space and improve retrieval speed. These index structures can quickly locate place name data sets that may belong to the same geographic area based on geographic coordinates, thus avoiding the enormous computational burden of comparing all place name data one by one.
[0065] After establishing a suitable spatial index, the next step is to calculate the distance between place name data. Euclidean distance is suitable for simple applications in a plane rectangular coordinate system, but when dealing with large-scale geographic areas, considering the influence of the earth's curvature, it is more accurate to use the Haversine formula to calculate the spherical distance between two points. This calculation method not only accurately reflects the actual distance between two locations, but also adapts to geographic ranges of different scales. For example, within a small area within a city, Euclidean distance may be sufficient; however, in a larger area across cities, the use of the Haversine formula becomes indispensable.
[0066] Next, we need to set a reasonable distance threshold. This threshold determines which place-name data pairs are considered geographically close and retained. If the threshold is set too low, some place-name data pairs that actually refer to the same entity but have slightly different geographic coordinates due to measurement errors or recording discrepancies may be missed. Conversely, if the threshold is too high, a large number of unrelated place-name data may be incorrectly classified into the same category, increasing the complexity of subsequent processing. Therefore, it is particularly important to carefully adjust and verify this threshold based on the specific data characteristics and application scenarios.
[0067] To further optimize the screening results, other auxiliary information can be incorporated for comprehensive judgment. For example, in addition to the proximity of geographic coordinates, the consistency of the place name's attribute information is also an important consideration. If two place names are not only geographically close but also have a certain degree of consistency in certain key attributes (such as administrative divisions and categories), then these place name data are more likely to refer to the same entity. This approach can further improve the accuracy of screening results without significantly increasing computational complexity.
[0068] In the above-mentioned method for automatic quality inspection of multi-source place name data, the step S3 extracts preliminary candidate pairs of place name data to be inspected from the set of preliminary candidate pairs of place name data, and the preliminary candidate pairs of place name data to be inspected include first place name data and second place name data. Specifically, since the preliminary candidate pairs of place name data obtained through geographical proximity screening may contain some entities with close coordinates but large name differences (such as "Chaoyang Park" and "Chaoyang District"), it is necessary to perform further pairing verification on each preliminary candidate pair of place name data to determine whether the two place name data contained therein belong to the same place name entity. Here, the present application extracts preliminary candidate pairs of place name data to be inspected from the set of preliminary candidate pairs of place name data, takes the preliminary candidate pairs of place name data to be inspected as an example, introduces semantic analysis technology based on deep learning, further explores the potential semantic association between the first place name data and the second place name data contained therein, thereby judging whether they point to the same place name entity, so as to realize automatic quality inspection of the preliminary candidate pairs of place name data.
[0069] In the above-mentioned multi-source place name data automatic quality inspection method, the step S4 is to calculate the entity semantic alignment between the first place name data and the second place name data. In a specific example of the present application, the step S4 includes: performing implicit semantic query space interactive alignment analysis on the first place name data and the second place name data to obtain the entity semantic alignment. Figure 3 Flowchart of sub-step S4 of the method for automatic quality inspection of multi-source place name data according to an embodiment of the present application. Figure 3 As shown, the step S4 includes the steps of: S41, performing structured mapping on the first place name data and the second place name data to obtain a first place name data structured mapping coding vector and a second place name data structured mapping coding vector; S42, performing semantic alignment analysis based on implicit query space sparse constraints on the first place name data structured mapping coding vector and the second place name data structured mapping coding vector to obtain a first-second place name data semantic interaction coding feature vector; S43, determining the entity semantic alignment degree based on the first-second place name data semantic interaction coding feature vector.
[0070] Specifically, the step S41 performs structured mapping on the first place name data and the second place name data to obtain a structured mapping encoding vector of the first place name data and a structured mapping encoding vector of the second place name data. In a specific example of the present application, the step S41 includes: using a semantic encoder based on the Bert model to semantically encode the first place name data and the second place name data to obtain a structured mapping encoding vector of the first place name data and a structured mapping encoding vector of the second place name data. It should be understood that since traditional text matching algorithms are highly dependent on the degree of overlap of characters on the surface of the text, they cannot capture the semantic equivalence of "Beijing" and "Beijing". Therefore, the present application introduces a semantic analysis technology based on deep learning, and uses the Bert model to semantically encode the first place name data and the second place name data to achieve structured mapping of the two. Specifically, the Bert model has powerful semantic understanding and representation capabilities through pre-training on a large-scale corpus. For place name data, which are words with specific cultural connotations and background knowledge, the Bert model can better capture the semantic connections between each word in the text data and encode its deep semantic features, thereby generating a first place name data structured mapping encoding vector and a second place name data structured mapping encoding vector containing contextual semantic information. In addition, by performing structured mapping on the first place name data and the second place name data, both are converted into vector representations in a high-dimensional semantic vector space, and then the distance or similarity between vectors can be used to measure the semantic relationship between place name data. For example, if "Beijing" and "Beijing" are close in the semantic vector space, they may point to the same place name entity. In this way, problems such as the same name in different places, different names for the same place, and semantic ambiguity can be effectively solved, thereby improving the accuracy and efficiency of place name data matching.
[0071] Specifically, the step S42 performs a semantic alignment analysis based on the sparse constraints of the implicit query space on the structured mapping coding vector of the first place name data and the structured mapping coding vector of the second place name data to obtain the semantic interaction coding feature vector of the first-second place name data. It should be understood that since the simple vector similarity calculation only relies on the global distance in the vector space, it lacks dynamic interaction analysis of fine-grained semantic units (such as administrative division levels, modifiers), and is disturbed by differences in word order and word segmentation granularity, and cannot accurately model cross-language alignment relationships. In this regard, the present application proposes a semantic alignment analysis method based on sparse constraints of implicit query space. By introducing the implicit query space, the model is allowed to dynamically search and match the relevant semantic features between the first place name data and the second place name data in a fine-grained manner in the semantic space, and explicitly model the potential association between the structured mapping coding vector of the first place name data and the structured mapping coding vector of the second place name data. At the same time, by introducing sparse constraints, the model is restricted to focus only on the most important semantic features, avoiding overfitting and noise interference, thereby improving the robustness and accuracy of the place name data matching alignment analysis. Among them, Figure 4 FIG4 is a flowchart of sub-step S42 of the method for automatic quality inspection of multi-source place name data according to an embodiment of the present application. Figure 4 As shown, the step S42 includes the steps of: S421, performing implicit query interaction based on the local semantic space on the first place name data structured mapping coding vector and the second place name data structured mapping coding vector to obtain a set of first-second place name data semantic interaction implicit query space joint coding matrices; S422, performing adaptive aggregation based on the implicit query space sparse constraint on the set of first-second place name data semantic interaction implicit query space joint coding matrices to obtain the first-second place name data semantic interaction coding feature vector.
[0072] Figure 5 Flowchart of sub-step S421 of the method for automatic quality inspection of multi-source place name data according to an embodiment of the present application. Figure 5 As shown, the step S421 includes the steps of: S4211, reconstructing the feature phase space of the first place name data structured mapping coding vector and the second place name data structured mapping coding vector based on one-dimensional convolution coding to obtain a set of first place name data local semantic feature coding vectors and a set of second place name data local semantic feature coding vectors; S4212, calculating the implicit query space joint coding matrix between any group of first place name data local semantic feature coding vectors and second place name data local semantic feature coding vectors in the set of the first place name data local semantic feature coding vectors and the set of the second place name data local semantic feature coding vectors to obtain a set of the first-second place name data semantic interaction implicit query space joint coding matrices.
[0073] In a specific example of the present application, step S4211 is expressed as follows:
[0074] ;
[0075] ;
[0076] in, A structured mapping encoding vector representing the first place name data, A structured mapping encoding vector representing the second place name data, is a one-dimensional convolutional coding network, represents the set of local semantic feature encoding vectors of the first place name data, 、 、 and Respectively represent the first, second, and third place name data in the set of local semantic feature encoding vectors of the first place name data and The first place name data local semantic feature encoding vector, is the number of vectors in the set of local semantic feature encoding vectors of the first place name data, represents the set of local semantic feature encoding vectors of the second place name data, 、 、 and Respectively represent the first, second, and third in the set of local semantic feature encoding vectors of the second place name data and The second local semantic feature encoding vector of place name data.
[0077] This application introduces the phase space reconstruction idea from the theory of dynamical systems, regards the first place name data structured mapping encoding vector and the second place name data structured mapping encoding vector as potential trajectories in the abstract feature space, and uses the sliding window mechanism of one-dimensional convolution to simulate the "delayed embedding" operation, and mines the semantic relevance and structural pattern implicit in the first place name data structured mapping encoding vector and the second place name data structured mapping encoding vector from different local perception perspectives. In this way, the original feature space is mapped to a higher-dimensional and more abstract semantic space through nonlinear feature transformation, which enhances the model's perception of the local semantic structure of the place name data and its robustness to noisy data or partial information loss. Even if some feature dimensions are interfered or missing, the stability of the semantic alignment analysis can still be guaranteed through the synergy of multiple local semantic features.
[0078] In a specific example of the present application, step S4212 is expressed as follows:
[0079] ;
[0080] in, represents the transpose of the matrix, represents the matrix multiplication operation, and They represent the first place name data semantic feature weight matrix and the second place name data semantic feature weight matrix, respectively. is the characteristic scale scaling factor, express and Joint encoding matrix of implicit query space between first-second place name data semantic interactions.
[0081] Specifically, by constructing an implicit query space, the local semantic feature coding vector of the first place name data and the local semantic feature coding vector of the second place name data are dynamically associated in the implicit query space, breaking through the limitations of explicit rules, automatically mining the potential semantic associations between place names, and forming a first-second place name data semantic interaction implicit query space joint coding matrix that can represent the complex interactive relationship between the two, thereby improving the model's generalization ability for semantic changes in place name data.
[0082] Figure 6 FIG4 is a flowchart of sub-step S422 of the method for automatic quality inspection of multi-source place name data according to an embodiment of the present application. Figure 6 As shown, the step S422 includes the steps of: S4221, performing redundant interaction topological reduction on each first-second place name data semantic interaction implicit query space joint coding matrix in the set of the first-second place name data semantic interaction implicit query space joint coding matrix to obtain a set of optimized first-second place name data semantic interaction implicit query space joint coding matrices; S4222, calculating the spatial sparse constraint factor of each optimized first-second place name data semantic interaction implicit query space joint coding matrix in the set of optimized first-second place name data semantic interaction implicit query space joint coding matrices to obtain a set of first-second place name data semantic interaction spatial sparse constraint factors; S4223, based on the set of the first-second place name data semantic interaction spatial sparse constraint factors, performing adaptive aggregation coding on the set of the optimized first-second place name data semantic interaction implicit query space joint coding matrices to obtain the first-second place name data semantic interaction coding feature vector.
[0083] In particular, in a preferred example of the present application, step S4221 includes: constructing a first place name data-second place name data semantic interaction phase matrix between the first place name data local semantic feature coding vector and the second place name data local semantic feature coding vector; based on the first place name data-second place name data semantic interaction phase matrix, optimizing and mapping the implicit query space joint coding matrix between the first place name data local semantic feature coding vector and the second place name data local semantic feature coding vector to obtain the optimized first-second place name data semantic interaction implicit query space joint coding matrix, which is expressed by the formula:
[0084] ;
[0085] ;
[0086] in, for The eigenvalues, for The eigenvalues, Indicates The exponential function operation with base , Represents the semantic interaction phase matrix between the first place name data and the second place name data, Indicates the semantic interaction phase matrix between the first place name data and the second place name data The eigenvalues of the position, represents the inverse matrix, Represents the optimized joint encoding matrix of the implicit query space of the semantic interaction between the first and second place name data.
[0087] Among them, considering that the joint encoding matrix of the implicit query space of the semantic interaction between the first and second place name data may contain a large number of redundant interactions (such as "Beijing" and "Beijing" repeatedly trigger the same semantic association under multiple convolution kernel perspectives), resulting in information overload. To this end, the present application introduces redundant interaction topological simplification, by performing local nonlinear phase representation analysis on each pair of local semantic feature coding vectors of the first place name data and the local semantic feature coding vectors of the second place name data, and constructing a phase matrix between the two. Then, drawing on the concept of trivial connection in differential geometry, the geometric structure of the joint encoding matrix of the implicit query space of the semantic interaction between the first and second place name data is flattened, and the redundant curvature is eliminated by the inverse transformation of the phase matrix, so as to achieve the projection dynamic simplification of the joint encoding matrix of the semantic interaction between the first and second place name data, so that the interaction relationship between the place name data presents linear decomposable characteristics in the latent space, which helps the subsequent analysis focus on the core discriminant features and improves the interpretability and computational efficiency of the feature interaction pattern in the semantic alignment calculation process.
[0088] In a specific example of the present application, step S4222 calculates the spatial sparse constraint factor of each optimized first-second place name data semantic interaction implicit query space joint encoding matrix in the set of optimized first-second place name data semantic interaction implicit query space joint encoding matrices to obtain a set of first-second place name data semantic interaction space sparse constraint factors. The first-second place name data semantic interaction space sparse constraint factor is the square of the F norm of the optimized first-second place name data semantic interaction implicit query space joint encoding matrix, which is expressed by the formula:
[0089] ;
[0090] in, is the correlation feature strength measurement function, It means to calculate the square of the Frobenius norm of the matrix. express The corresponding sparse constraint factor of the semantic interaction space of the first-second place name data.
[0091] Specifically, by converting the sparsity characteristics of implicit semantic interactions into a computable sparse constraint factor of the first-second place name data semantic interaction space, weight guidance is provided for the adaptive aggregation step. The size of the sparse constraint factor of the first-second place name data semantic interaction space directly reflects the density and intensity of effective semantic interactions in the optimized first-second place name data semantic interaction implicit query space joint encoding matrix, so that the model can automatically adjust the weight ratio of different semantic interaction relationships according to the sparse constraint factor of the first-second place name data semantic interaction space during subsequent alignment calculations. It should be understandable that the F-norm squared, as a calculation method for the sum of squares of the elements of the optimized joint encoding matrix of the implicit query space of the semantic interaction between the first and second place name data, can effectively characterize the overall energy distribution characteristics of the matrix elements. Especially when only a few elements in the optimized joint encoding matrix of the implicit query space of the semantic interaction between the first and second place name data are significant and most of them are close to zero, the F-norm can highlight the importance of key features in the sparse interaction pattern by amplifying the numerical contribution of significant elements, thereby accurately capturing the potential and non-explicit semantic associations between place name data, suppressing the interference of redundant or noise information, and thus improving the discrimination and robustness of the entity semantic alignment calculation.
[0092] In a specific example of the present application, step S4223 is expressed as follows:
[0093] ;
[0094] ;
[0095] in, represents the normalized exponential function, express The corresponding normalized first-second place name data semantic interaction space sparse constraint factor, Represents the sparsity fusion matrix of the implicit coding features of the semantic interaction between the first and second place name data, represents the feature shape reshaping, Represents the semantic interaction encoding feature vector of the first-second place name data.
[0096] Specifically, since the optimized joint encoding matrices of the implicit query space of the first-second place-name data semantic interaction capture interactive information from different local feature perspectives, each matrix may contain redundant noise information or complementary effective features. Directly integrating them by averaging or concatenating them will dilute key semantic associations and amplify secondary or erroneous information. Therefore, the quality and reliability of the interactive information from each local perspective are dynamically evaluated using the sparse constraint factor of the first-second place-name data semantic interaction space. Optimized joint encoding matrices of the implicit query space of the first-second place-name data semantic interaction space with higher sparse constraint factors tend to correspond to more focused and significant semantic association patterns, while matrices with lower factors may contain redundant or low-value information. Through this adaptive aggregation coding approach, the importance of multi-perspective interactive information is ranked and filtered, allowing optimized joint encoding matrices of the implicit query space of the first-second place-name data semantic interaction space with higher sparse constraint factors to dominate the fusion process, thereby suppressing noise interference, preserving complementary semantic features across perspectives, and enhancing the targeted nature of feature fusion.
[0097] Specifically, the step S43 determines the entity semantic alignment based on the semantic interaction encoding feature vector of the first and second place name data. That is, in order to convert the semantic interaction encoding feature vector of the first and second place name data into an interpretable matching probability to support relevant decision-making, the present application adopts a decoding algorithm, such as a regression model based on a neural network, by taking the semantic interaction encoding feature vector of the first and second place name data as input, and using the trained regression model to perform feature decoding on the semantic association interaction pattern between the place name data implicit in the vector, so as to map the high-dimensional semantic interaction feature into a scalar alignment, and use the Sigmoid function to output an entity semantic alignment value between 0 and 1. The closer the value is to 1, the higher the probability that the first place name data and the second place name data point to the same place name entity, which helps to intuitively judge the degree of semantic similarity between the first place name data and the second place name data.
[0098] In the above-mentioned automated quality inspection method for multi-source place name data, step S5 determines whether the first place name data and the second place name data belong to the same place name entity based on the comparison between the entity semantic alignment and the preset threshold. Specifically, if the entity semantic alignment is greater than the preset threshold (for example, 0.8), it is determined that the first place name data and the second place name data point to the same place name entity, thereby confirming that the pair of place name data is a valid match. Conversely, if the entity semantic alignment is lower than the preset threshold, it is determined that the pair of place name data does not point to the same place name entity, is an invalid match, and is eliminated from the candidate pairs. In this way, place name data pairs that truly point to the same place name entity can be further screened out, thereby improving the accuracy and reliability of place name data matching, and providing strong support for the integration, management and application of place name data.
[0099] In summary, an automated quality inspection method for multi-source place name data based on an embodiment of the present application is illustrated, which first accesses place name data from multiple data sources, and performs preliminary matching screening on the accessed place name data based on geographical proximity. Then, a semantic analysis technology based on deep learning is further introduced, and semantic parsing and implicit semantic query interaction analysis are performed on the two place name data in the candidate pair to mine the potential semantic association between the two, determine the entity semantic alignment between the place name data, and then determine whether the two belong to the same place name entity based on the comparison of the entity semantic alignment with a preset threshold. In this way, problems such as the same name in different places, different names for the same place, and semantic ambiguity can be effectively solved, the quality and integration efficiency of place name data can be improved, and accurate place name data support can be provided for applications such as urban construction planning.
[0100] Furthermore, an automated quality inspection system for multi-source place name data is also provided.
[0101] Figure 7 FIG is a block diagram of an automatic quality inspection system for multi-source place name data according to an embodiment of the present application. Figure 7As shown, the multi-source place name data automatic quality inspection system 100 according to the embodiment of the present application includes: a data access and integration module 110, which is used to access place name data from multiple data sources to obtain a set of place name data, wherein the place name data includes place name names, coordinates and attribute information; a data candidate pair screening module 120, which is used to screen the set of place name data for preliminary candidate pairs based on geographical location proximity to obtain a set of preliminary candidate pairs of place name data; a candidate pair extraction module 130, which is used to extract preliminary candidate pairs of place name data to be inspected from the set of preliminary candidate pairs of place name data, wherein the preliminary candidate pairs of place name data to be inspected are selected from the set of preliminary candidate pairs of place name data. The pair includes first place name data and second place name data; an entity semantic alignment analysis module 140 is used to calculate the entity semantic alignment between the first place name data and the second place name data, wherein calculating the entity semantic alignment between the first place name data and the second place name data includes: performing implicit semantic query space interactive alignment analysis on the first place name data and the second place name data to obtain the entity semantic alignment; a same-name entity determination module 150 is used to determine whether the first place name data and the second place name data belong to the same place name entity based on a comparison between the entity semantic alignment and a preset threshold.
[0102] The specific operations of each module in the above multi-source place name data automatic quality inspection system have been referenced above. Figures 1 to 6 The method for automatic quality inspection of multi-source place name data has been introduced in detail, and therefore, its repeated description will be omitted.
[0103] The basic principles of the present invention have been described above in conjunction with specific embodiments. However, it should be noted that the advantages, strengths, and effects mentioned in the present invention are merely illustrative and non-limiting, and should not be construed as necessarily possessed by each embodiment of the present invention. Furthermore, the specific details of the above embodiments are provided for illustrative purposes and to facilitate understanding, and are not intended to be limiting. These details do not necessarily limit the present invention to being implemented using these specific details.
[0104] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, please refer to the relevant description of other embodiments. In the several embodiments provided by the present invention, it should be understood that the disclosed system and method can be implemented in other ways. For example, the system embodiment described above is only schematic. For example, the unit division is only a logical function division, and there may be other division methods in actual implementation. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.
[0105] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be embodied in other specific forms without departing from the spirit or essential characteristics of the invention. Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims, not the foregoing description, and all variations within the meaning and range of equivalents of the claims are intended to be encompassed therein. Any reference to a figure in a claim should not be construed as limiting the claim to which it relates.
[0106] In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units stated in the system claims can also be implemented by one unit through software or hardware.
[0107] Finally, it should be noted that the above description has been provided for purposes of illustration and description. Furthermore, the above embodiments are intended only to illustrate the technical solutions of the present invention and are not intended to be limiting. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art will appreciate that the technical solutions of the present invention may be modified or replaced with equivalents without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. An automated quality inspection method for multi-source place name data, characterized in that: include: Accessing place name data from multiple data sources to obtain a set of place name data, wherein the place name data includes place name, coordinates and attribute information; Based on geographical location proximity, screening the set of place name data for preliminary candidate pairs to obtain a set of preliminary candidate pairs of place name data; Extracting preliminary candidate pairs of place name data to be quality checked from the set of preliminary candidate pairs of place name data, wherein the preliminary candidate pairs of place name data to be quality checked include first place name data and second place name data; Calculating an entity semantic alignment between the first place name data and the second place name data, wherein calculating the entity semantic alignment between the first place name data and the second place name data includes: performing a latent semantic query space interactive alignment analysis on the first place name data and the second place name data to obtain the entity semantic alignment; Determining whether the first place name data and the second place name data belong to the same place name entity based on a comparison between the entity semantic alignment and a preset threshold; Performing a latent semantic query space interactive alignment analysis on the first place name data and the second place name data to obtain the entity semantic alignment, including: Performing structured mapping on the first place name data and the second place name data to obtain a first place name data structured mapping encoding vector and a second place name data structured mapping encoding vector; Performing implicit query interaction based on a local semantic space on the first place name data structured mapping code vector and the second place name data structured mapping code vector to obtain a set of first-second place name data semantic interaction implicit query space joint coding matrices; Constructing a first place name data-second place name data semantic interaction phase matrix between the first place name data local semantic feature encoding vector and the second place name data local semantic feature encoding vector; Based on the first place name data-second place name data semantic interaction phase matrix, the implicit query space joint coding matrix between the local semantic feature coding vector of the first place name data and the local semantic feature coding vector of the second place name data is optimized and mapped to obtain the optimized first-second place name data semantic interaction implicit query space joint coding matrix, which is expressed as follows: ; ; in, for The eigenvalues, Indicates the The first place name data local semantic feature encoding vector, for The eigenvalues, No. The second place name data local semantic feature encoding vector, Represents the exponential function operation with e as the base, Represents the semantic interaction phase matrix between the first place name data and the second place name data, Indicates the semantic interaction phase matrix between the first place name data and the second place name data The eigenvalues of the position, represents the inverse matrix, represents the joint encoding matrix of the implicit query space of the semantic interaction between the first and second place name data, Represents the optimized joint encoding matrix of the implicit query space of the semantic interaction between the first and second place name data; Calculate the spatial sparse constraint factor of each optimized first-second place name data semantic interaction implicit query space joint encoding matrix in the set of the optimized first-second place name data semantic interaction implicit query space joint encoding matrix to obtain a set of first-second place name data semantic interaction space sparse constraint factors, wherein the first-second place name data semantic interaction space sparse constraint factor is the square of the F norm of the optimized first-second place name data semantic interaction implicit query space joint encoding matrix, and is expressed by the formula: ; in, is the correlation feature strength measurement function, It means to calculate the square of the Frobenius norm of the matrix. express The corresponding sparse constraint factor of the semantic interaction space of the first-second place name data; Based on the set of sparse constraint factors of the first-second place name data semantic interaction space, adaptively aggregate coding is performed on the set of optimized first-second place name data semantic interaction implicit query space joint coding matrices to obtain a first-second place name data semantic interaction coding feature vector; The entity semantic alignment is determined based on the semantic interaction encoding feature vector of the first and second place name data.
2. The method for automatic quality inspection of multi-source place name data according to claim 1, characterized in that: Performing structured mapping on the first place name data and the second place name data to obtain a first place name data structured mapping encoding vector and a second place name data structured mapping encoding vector, further comprising: The first place name data and the second place name data are semantically encoded using a semantic encoder based on the Bert model to obtain a structured mapping encoding vector of the first place name data and a structured mapping encoding vector of the second place name data.
3. The method for automatic quality inspection of multi-source place name data according to claim 2, characterized in that: Performing implicit query interaction based on a local semantic space on the first place name data structured mapping code vector and the second place name data structured mapping code vector to obtain a set of first-second place name data semantic interaction implicit query space joint coding matrices, including: Performing feature phase space reconstruction based on one-dimensional convolution coding on the first place name data structured mapping code vector and the second place name data structured mapping code vector to obtain a set of first place name data local semantic feature code vectors and a set of second place name data local semantic feature code vectors; Calculate the implicit query space joint coding matrix between any set of the first place name data local semantic feature coding vectors and the second place name data local semantic feature coding vectors in the set of the first place name data local semantic feature coding vectors and the set of the second place name data local semantic feature coding vectors to obtain the set of the first-second place name data semantic interaction implicit query space joint coding matrices.
4. The method for automatic quality inspection of multi-source place name data according to claim 3, characterized in that: Determining the entity semantic alignment based on the semantic interaction encoding feature vector of the first and second place name data includes: The semantic interaction encoding feature vector of the first and second place name data is input into a decoder-based semantic alignment quality inspection module to obtain the entity semantic alignment degree.
5. An automated quality inspection system for multi-source place name data, used to execute the automated quality inspection method for multi-source place name data according to any one of claims 1 to 4, characterized in that: include: A data access and integration module, configured to access place name data from multiple data sources to obtain a collection of place name data, wherein the place name data includes place name, coordinates, and attribute information; a data candidate pair screening module, configured to screen the set of place name data for preliminary candidate pairs based on geographical location proximity to obtain a set of preliminary candidate pairs of place name data; a candidate pair extraction module for extracting preliminary candidate pairs of place name data to be inspected from the set of preliminary candidate pairs of place name data, wherein the preliminary candidate pairs of place name data to be inspected include first place name data and second place name data; an entity semantic alignment analysis module, configured to calculate an entity semantic alignment degree between the first place name data and the second place name data, wherein calculating the entity semantic alignment degree between the first place name data and the second place name data comprises: performing a latent semantic query space interactive alignment analysis on the first place name data and the second place name data to obtain the entity semantic alignment degree; The homonymous entity determination module is used to determine whether the first place name data and the second place name data belong to the same place name entity based on a comparison between the entity semantic alignment and a preset threshold.
Citation Information
Patent Citations
Multi-relation perception heterogeneous graph visual question and answer method fusing syntax tree
CN117034185A
Multi-modal medical image quality inspection system based on deep learning
CN118334036A
Entity alignment method and system based on deentanglement graph neural network
CN119250065A