Hierarchical structured address-based high-precision duplicate removal fusion library building method and system

By using dictionary-based and deep learning methods to perform hierarchical and structured processing of Shanghai addresses, integrating authoritative data and correcting conflicts, the problem of high duplication rate of Shanghai address data is solved, achieving high-precision address fusion and data quality improvement, supporting smart logistics and O2O services.

CN121958286APending Publication Date: 2026-05-01上海市大数据中心
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
上海市大数据中心
Filing Date
2026-01-06
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Shanghai's address data suffers from contradictions in its dual-track system, confusion regarding administrative boundaries, significant differences in urban and rural address structures, and lagging dynamic updates, resulting in a high rate of address data duplication. This hinders the high-precision application of smart logistics, urban management, and O2O services.

Method used

Initial word segmentation is performed using a dictionary-based algorithm cluster and rule-based segmentation marker method. Sequence labeling is performed using a deep learning model. Authoritative data is integrated and conflict detection and forced correction are implemented. Anchor relationships are defined, and field information is integrated based on the principle of prioritizing official sources. The information is stored in a dedicated database and updated regularly. The address database is also optimized.

Benefits of technology

It has achieved high-precision fusion of Shanghai address data, reduced duplication rate, improved data quality, and supported high-precision applications in smart logistics, urban management, and O2O services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121958286A_ABST
    Figure CN121958286A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of deep learning, in particular to a hierarchical structured address-based high-precision duplicate removal fusion library building method, which comprises the following steps of: carrying out preliminary word segmentation and preprocessing on an address text by adopting a dictionary-based algorithm cluster and rule segmentation marking method, and carrying out conflict detection, authoritative verification and forced correction to obtain a hierarchical structured address-based duplicate removal fusion library; for old-to-new-code adaptation, associating old codes to new codes, defining an anchor point relationship between urban areas and suburbs, integrating field information by adopting an official source priority and non-NULL supplement principle, forming a complete address, storing processed address data into a special database, constructing a complete Shanghai address database, and updating an authoritative data layer regularly; and according to the new change and the business demand optimization model and algorithm, obtaining a continuously optimized address library. According to the method, the problems of high repetition rate and low fusion precision of address data caused by'double-track system expression, administrative division confusion, urban and rural structure difference and dynamic update lag 'are solved.
Need to check novelty before this filing date? Find Prior Art

Description

A High-Precision Deduplication and Fusion Database Construction Method and System Based on Hierarchical Structured Addresses Technical Field

[0001] This invention belongs to the field of deep learning technology, specifically relating to a high-precision deduplication and fusion library construction method and system based on hierarchical structured addresses. Background Technology

[0002] Shanghai's address system exhibits characteristics of "multi-level, multi-form, and highly dynamic," and existing address data fusion schemes have significant adaptation defects. Specific problems include: 1. A prominent dual-track contradiction in address representation: Shanghai addresses commonly exhibit a coexistence of "official administrative codes" and "common commercial / residential names." For example, "No. 18, Zhongshan East Road, Bund Street, Huangpu District" is the official standard address, while the corresponding commercial name "Bund 18" is frequently used in scenarios such as logistics orders and food delivery; the official code for "Lane 18, Xuhong North Road, Xujiahui Street, Xuhui District" cannot be linked to the commonly used name "Xujiahui Jingyuan" by residents, resulting in the same address being recorded repeatedly.

[0003] 2. Confusion regarding administrative boundaries: Some areas in Shanghai have "enclaves" or "interspersed areas," and the names of roads in adjacent districts can easily lead to misclassification. For example, "Dalian Road," which borders Hongkou District and Yangpu District, is sometimes mistakenly classified as "No. 1500 Dalian Road" (which actually belongs to Yangpu District) in Hongkou District by some logistics data; "Caobao Road," which borders Minhang District and Xuhui District, often has "No. 88 Caobao Road" (which belongs to Xuhui District) marked as Minhang District, directly causing cross-district address deduplication failures.

[0004] 3. Significant differences in urban and rural address structures: Addresses in Shanghai's central urban areas (such as Huangpu and Jing'an districts) are mainly based on a refined structure of "street-lane-building-room," while suburban areas (such as Chongming and Jinshan districts) have a large number of non-standardized addresses based on "natural village-village number." For example, "No. 35, Group 12, Yu'an Village, Chenjia Town, Chongming District" lacks a "lane number" hierarchy similar to that in urban areas, and the existing unified modeling scheme cannot be adapted, resulting in insufficient accuracy in suburban address fusion.

[0005] 4. Delayed Address Updates: Shanghai's rapid urban renewal often leads to issues such as "address names remaining unchanged but administrative affiliations adjusted" and "newly added buildings lacking codes" in redeveloped communities and newly built commercial areas. For example, after the redevelopment of "Zhongyuan Liangwan City" in Putuo District, "Lane 1168, Yuanjing Road" was added, but some data sources still use the old code "Lane 2138, Zhongshan North Road"; in Pudong New Area's "Qiantan Business District," the address codes of newly built buildings are not synchronized between postal and map data, resulting in data conflicts.

[0006] The aforementioned issues have resulted in a duplication rate of over 35% in Shanghai's multi-source address data (based on a sample of 5 million address data entries from a logistics company in Shanghai). This low data quality severely hinders the implementation of high-precision applications in Shanghai, such as smart logistics (e.g., precise delivery for the "last mile"), urban management (e.g., grid-based governance), and O2O services (e.g., instant retail delivery). Summary of the Invention

[0007] In response to the above situation, this invention provides a high-precision deduplication and fusion database construction method and system based on hierarchical structured addresses, which can solve the problems of high duplication rate and low fusion accuracy caused by address data due to "dual-track expression, confusion of administrative divisions, differences in urban and rural structures, and lag in dynamic updates". To achieve the above objectives, the present invention adopts the following technical solution: The high-precision deduplication and fusion database construction method based on hierarchical structured addresses includes the following steps: using dictionary-based algorithm clustering and rule-based segmentation flag method to perform preliminary word segmentation and preprocessing on address text, and using a deep learning model for sequence labeling to obtain 18-level structured results; integrating authoritative data, maintaining the authoritative data layer, and through conflict detection, authoritative verification, and forced correction, old codes are associated with new codes to unify administrative affiliation, resulting in Shanghai address records; defining the anchor point relationship between urban and suburban areas, performing administrative layer pre-matching and core anchor point matching, and obtaining Shanghai address fusion results through high-dimensional vector clustering deduplication, BERT semantic similarity calculation, and fusion judgment formula; integrating field information according to the principle of official source priority and non-NULL supplementation to form a complete address, recording the anchor point mapping relationship in the address database association table, and storing aliases and old codes in specific fields to obtain Shanghai addresses and anchor point mapping relationships; storing the processed address data in a dedicated database to construct a complete Shanghai address database, regularly updating the authoritative data layer, and optimizing the model and algorithm according to new changes and business needs to obtain a continuously optimized address database.

[0008] Furthermore, the integration and maintenance of authoritative data, through conflict detection, authoritative verification, and forced correction, links old codes to new codes for urban renewal adaptation, unifying administrative affiliation and obtaining Shanghai address records, includes the following steps: Integrating localized authoritative data in Shanghai using a combination of periodic collection and proactive reporting. This authoritative data includes road names and districts, place names and districts, and urban renewal address update information; the authoritative data layer is updated and maintained in real time; conflicts are detected by comparing low-level and high-level information of different address records, and verification is performed using accurate mapping relationships in the authoritative data layer to extract erroneous fields; erroneous fields are forcibly corrected to correct content, unifying administrative affiliation. For urban renewal addresses, the new codes associated with the old codes are extracted by querying the urban renewal address update table, resulting in unified and accurate Shanghai address data with unified administrative affiliation.

[0009] Furthermore, the process of defining the anchor point relationship between urban and suburban areas, performing administrative-level pre-matching and core anchor point matching, and obtaining the Shanghai address fusion result through high-dimensional vector clustering for deduplication, BERT semantic similarity calculation, and a fusion judgment formula includes the following steps: defining the anchor point relationship in Shanghai's urban areas using lane numbers as official codes and courtyards as colloquial terms, and in suburban areas using village numbers as official codes and natural villages as colloquial terms; employing administrative-level pre-matching, prioritizing the verification of the first four levels of address fields and correcting them to standard administrative affiliation; performing core anchor point matching, extracting address entity features through high-dimensional vector clustering for deduplication, and calculating the semantic similarity of core anchor point fields using the BERT Siamese network model; combining the fusion judgment formula to obtain the total deduplication score, and then obtaining the address's latitude and longitude through geographic decoding, calculating the distance, and determining those that meet the distance condition as having overlapping physical locations, thus obtaining the accurately fused Shanghai address result.

[0010] Furthermore, the method of integrating field information using the principle of prioritizing official sources and supplementing with non-NULL values ​​to form a complete address, recording anchor mapping relationships in the address database association table, and storing aliases and old codes in specific fields to obtain Shanghai addresses and anchor mapping relationships includes the following steps: Using the principle of prioritizing official Shanghai sources and supplementing with non-NULL values, field information from different data sources is integrated, prioritizing field content from official Shanghai data sources; when official data is empty, supplementation is extracted from non-empty data sources to form a complete address; by sorting out the official codes and colloquial information in the address, the anchor mapping relationships are recorded in detail in the address database association table; aliases and old codes unique to Shanghai addresses are extracted and stored in specific fields, thereby improving the address recognition rate in local scenarios.

[0011] Furthermore, the process of storing the processed address data in a dedicated database to construct a complete Shanghai address database, regularly updating the authoritative data layer, and optimizing models and algorithms based on new changes and business needs to obtain a continuously optimized address database includes the following steps: Using a database management module, the Shanghai address data, after word segmentation, structuring, conflict correction, and fusion, is accurately stored in a specially constructed database to build a complete Shanghai address database; By setting scheduled tasks and a real-time monitoring mechanism, the latest data is regularly obtained from authoritative channels to update the authoritative data layer; Based on newly emerging address changes and constantly evolving business needs, targeted optimizations are performed, continuously adjusting parameters and strategies to obtain Shanghai address deduplication and fusion results with continuously improving accuracy and effectiveness.

[0012] Furthermore, the formula for determining fusion is as follows: Where S_total is the total deduplication score calculated by the system, and W_vector is the similarity weight of the high-dimensional vector. Fusion is only performed when S_total exceeds the threshold T_fusion.

[0013] The second aspect of this invention provides a high-precision deduplication and fusion database construction system based on hierarchical structured addresses. This system includes the following modules: a structured module, used to perform preliminary word segmentation and preprocessing of address text using a dictionary-based algorithm clustering and rule-based segmentation marker method, and to perform sequence labeling using a deep learning model to obtain 18-level structured results; a conflict correction module, used to integrate authoritative data, maintain the authoritative data layer, and through conflict detection, authoritative verification, and forced correction, associate old codes with new codes for old-to-new adaptation, unifying administrative affiliation to obtain Shanghai address records; and a dual anchoring module, used to define the relationship between urban and suburban anchor points and perform pre-matching at the administrative level. The system employs a core anchor point matching mechanism, employing high-dimensional vector clustering for deduplication, BERT semantic similarity calculation, and a fusion determination formula to obtain the Shanghai address fusion result. The relationship mapping module integrates field information using a principle of prioritizing official sources and supplementing with non-NULL values ​​to form a complete address. Anchor point mapping relationships are recorded in the address database association table, and aliases and old codes are stored in specific fields to obtain the Shanghai address and its anchor point mapping relationship. The continuous optimization module stores the processed address data in a dedicated database to build a complete Shanghai address database. It regularly updates the authoritative data layer and optimizes the model and algorithm based on new changes and business needs to obtain a continuously optimized address database.

[0014] A third aspect of the present invention provides a high-precision deduplication and fusion database building device based on hierarchical structured addresses. The high-precision deduplication and fusion database building device based on hierarchical structured addresses includes a memory and at least one processor. The memory stores instructions. The at least one processor invokes the instructions in the memory to cause the high-precision deduplication and fusion database building device based on hierarchical structured addresses to perform the steps of the high-precision deduplication and fusion database building method based on hierarchical structured addresses as described in any of the preceding claims.

[0015] A fourth aspect of the present invention provides a computer-readable storage medium storing instructions, characterized in that, when executed by a processor, the instructions implement the steps of the high-precision deduplication and fusion library construction method based on hierarchical structured addresses as described in any one of the preceding claims.

[0016] In the technical solution provided by this invention, a dictionary-based algorithm cluster and rule-based segmentation method are used to perform preliminary word segmentation and preprocessing on address text. A deep learning model is used for sequence labeling to obtain 18-level structured results. Authoritative data is integrated and an authoritative data layer is maintained. Through conflict detection, authoritative verification, and forced correction, old codes are associated with new codes for old-to-new adaptation, and administrative affiliation is unified to obtain Shanghai address records. The anchor point relationship between urban and suburban areas is defined, and administrative-level pre-matching and core anchor point matching are performed. Through high-dimensional vector clustering for deduplication, BERT semantic similarity calculation, and fusion judgment formula, the Shanghai address fusion result is obtained. Field information is integrated according to the principle of official source priority and non-NULL supplementation to form a complete address. Anchor point mapping relationship is recorded in the address database association table. Aliases and old codes are stored in specific fields to obtain Shanghai addresses and anchor point mapping relationships. The processed address data is stored in a dedicated database to build a complete Shanghai address database. The authoritative data layer is updated regularly, and the model and algorithm are optimized according to new changes and business needs to obtain a continuously optimized address database. This invention solves the problems of high repetition rate and low fusion accuracy of address data caused by "dual-track expression, confusion of administrative divisions, differences in urban and rural structures, and lag in dynamic updates". Attached Figure Description

[0017] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention.

[0018] Figure 1 is a schematic diagram of the first embodiment of the high-precision deduplication and fusion database construction method based on hierarchical structured addresses in this invention.

[0019] Figure 2 is a schematic diagram of the second embodiment of the high-precision deduplication and fusion database construction method based on hierarchical structured addresses in this invention.

[0020] Figure 3 is a schematic diagram of the third embodiment of the high-precision deduplication and fusion database construction method based on hierarchical structured addresses in this invention.

[0021] Figure 4 is a schematic diagram of the fourth embodiment of the high-precision deduplication and fusion database construction method based on hierarchical structured addresses in this invention.

[0022] Figure 5 is a schematic diagram of the fifth embodiment of the high-precision deduplication and fusion database construction method based on hierarchical structured addresses in this invention. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0024] Those skilled in the art will understand that, unless specifically stated otherwise, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the term “comprising” as used in this specification means the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0025] The high-precision deduplication and fusion database construction method based on hierarchical structured addresses, as shown in Figure 1, includes the following steps: First, a dictionary-based algorithm clustering and rule-based segmentation flag method are used to perform preliminary word segmentation and preprocessing of the address text. A deep learning model is used for sequence labeling to obtain an 18-level structured result. Second, authoritative data is integrated, and an authoritative data layer is maintained. Through conflict detection, authoritative verification, and forced correction, old codes are associated with new codes to unify administrative affiliation, resulting in Shanghai address records. Third, the anchor point relationship between urban and suburban areas is defined, and administrative layer pre-matching and core anchor point matching are performed. High-dimensional vector clustering for deduplication, BERT semantic similarity calculation, and fusion judgment formulas are used to obtain the Shanghai address fusion result. Fourth, field information is integrated using the principle of prioritizing official sources and supplementing with non-NULL values ​​to form a complete address. Anchor point mapping relationships are recorded in the address database association table, and aliases and old codes are stored in specific fields to obtain Shanghai addresses and anchor point mapping relationships. Fifth, the processed address data is stored in a dedicated database to construct a complete Shanghai address database. The authoritative data layer is updated regularly, and the model and algorithm are optimized based on new changes and business needs to obtain a continuously optimized address database.

[0026] As shown in Figure 2, in this embodiment, dictionary-based forward and reverse maximum matching, optimal matching, and rule-based segmentation labeling methods are used to perform preliminary word segmentation and preprocessing on the input unstructured Shanghai address text. Through Bi-LSTM-CRF and ID-CNN models in deep learning, the segmented sequence is accurately labeled according to the 18-level label system to extract key information. Domain adaptive learning is used to transfer the authoritative source resolution capability to low-quality data. Multiple models are integrated, and through a high-confidence priority and voting mechanism, the 18-level structured result of the Shanghai address is obtained.

[0027] Targeting the address characteristics of Shanghai, a multi-strategy integrated address structuring and domain-adaptive engine was designed to decompose unstructured / semi-structured address data into 18 hierarchical fields.

[0028] The key focus is on optimizing the address hierarchy definition unique to Shanghai:

[0029] For Shanghai, a customized 18-level address structure is used. This engine achieves accurate and adaptive structuring of address text from all data sources. The specific technologies used include: 1.1 Basic word segmentation and preprocessing: Dictionary-based algorithm clusters (forward / reverse maximum matching method, best matching method) and rule-based segmentation marker method are used as the robust baseline for address parsing.

[0030] 1.2 Deep Learning Named Entity Recognition (NER): A Bi-LSTM-CRF / ID-CNN model is used to perform sequence labeling of the input sequence using an 18-level labeling system. The model training set focuses on optimizing the accurate extraction of ambiguous boundary fields such as ^9 (lane number) and ^13 (courtyard).

[0031] 1.3 Domain Adaptive Learning: A domain adaptive model is trained using a GAN architecture. Adversarial training is performed using an address parsing generator and a domain discriminator, enabling the model to transfer the parsing capabilities of authoritative sources to low-quality user input data without requiring the labeling of new data, thus achieving adaptive parsing for different business data sources.

[0032] 1.4 Model Integration and Decision Making: The system integrates and runs multiple models, and finally generates the final 18-level structured result through a high-confidence priority mechanism and a voting mechanism.

[0033] As shown in Figure 3, this embodiment adopts a combination of periodic collection and proactive reporting to integrate authoritative local data in Shanghai. The authoritative data includes road names and districts, place names and districts, and urban renewal address update information. The authoritative data layer is updated and maintained in real time. By comparing the low-level and high-level information of different address records, conflicts are detected. Verification is performed using the accurate mapping relationship in the authoritative data layer to extract erroneous fields. Erroneous fields are forcibly corrected to correct content to unify administrative affiliation. For urban renewal addresses, the new codes associated with the old codes are extracted by querying the urban renewal address update table, resulting in unified and accurate Shanghai address data with unified administrative affiliation.

[0034] As shown in Figure 4, in this embodiment, anchor relationships are defined in Shanghai's urban areas using lane numbers as official codes and courtyards as colloquial terms, and in suburban areas using village numbers as official codes and natural villages as colloquial terms. Administrative-level pre-matching is employed, prioritizing the verification of the first four levels of address fields and correcting them to standard administrative affiliation. Core anchor matching is then performed, extracting address entity features through high-dimensional vector clustering for deduplication. The semantic similarity of core anchor field fields is calculated using a BERT twin network model. A total deduplication score is derived by combining the fusion judgment formula. Then, the address's latitude and longitude are obtained through geographic decoding, and the distance is calculated. Addresses meeting the distance condition are considered to have overlapping physical locations, resulting in an accurately fused Shanghai address result.

[0035] As shown in Figure 5, in this embodiment, the principle of prioritizing official Shanghai data sources and supplementing with non-NULL data is adopted. Field information from different data sources is integrated, and field content from official Shanghai data sources is selected first. When official data is empty, supplementary data is extracted from non-empty data sources to form a complete address. By sorting out the official codes and common names in the address, the anchor mapping relationship is recorded in detail in the address database association table. The unique aliases and old codes of Shanghai addresses are extracted and stored in specific fields to improve the address recognition rate in local scenarios.

[0036] In this embodiment, a database management module is used to accurately store the Shanghai address data, which has undergone word segmentation, structuring, conflict correction, and fusion, into a specially constructed database, thus building a complete Shanghai address database. By setting scheduled tasks and a real-time monitoring mechanism, the latest data is regularly obtained from authoritative channels, and the authoritative data layer is updated. Based on newly emerging address changes and constantly changing business needs, targeted optimizations are made, and parameters and strategies are continuously adjusted to obtain Shanghai address deduplication and fusion results with continuously improving accuracy and effectiveness.

[0037] To address the common issues of address confusion regarding district boundaries and mislabeling across districts in Shanghai, a process for correcting administrative division conflicts is constructed, centered on "authoritative geographical data of Shanghai," to ensure the uniqueness of address administrative affiliation.

[0038] A. The Shanghai Authoritative Data Layer (A_data_sh) maintains and integrates Shanghai's local authoritative data in real time, specifically including: Shanghai Road Name-District Mapping Table: containing the precise district affiliation of 2000+ major roads in Shanghai, such as clearly stating that "Dalian Road (Huangxing Road-Yangshupu Road section) belongs to Yangpu District" and "Caobao Road (Guilin Road-Hongmei Road section) belongs to Xuhui District", resolving the ambiguity of road affiliation across districts; Shanghai Place Name-District Mapping Table: covering the unique district affiliation of 5000+ residential areas, business districts, and buildings in Shanghai, such as "Xujiahui Jingyuan" uniquely belonging to Xujiahui Subdistrict of Xuhui District, and "Qiantan Oriental Banyan Tree" uniquely belonging to Qiantan Subdistrict of Pudong New Area; Shanghai Urban Renewal Address Update Table: synchronizing with the Shanghai Municipal Commission of Housing and Urban-Rural Development's urban renewal project data, recording the correspondence between old codes and new addresses, such as "Putuo District Zhongyuan Liangwan City old code: Lane 2138, Zhongshan North Road → new code: Lane 1168, Yuanjing Road".

[0039] B. Shanghai Customized Forced Correction Algorithm: When two Shanghai address records exhibit "low-level consistency but high-level (^2 zone) conflict," the algorithm first addresses administrative conflict correction by performing a vertical normalization check using an authoritative geographic database (graph database verification of parent-child relationships) to forcibly correct errors in the ^2 field. Next, a horizontal address check (spatial discrimination) is performed using a GIS boundary layer and an SVM / random forest model to verify whether the address coordinates legally fall within the administrative boundaries corresponding to the ^2 field, thus determining the conflicting records. Finally, address ambiguity resolution is performed by training a multi-classifier to determine the geographic entity category (residential, commercial, industrial, etc.) of the address record based on the ^8, ^9, and ^13 fields. If two ^8, ^9, and ^13 fields are similar but the entity categories are different, merging is immediately prohibited. Perform the following steps: 1. Conflict Detection: For example, record 1 is "Lane 18, Xuhong North Road, Xuhui District", and record 2 is "Lane 18, Xuhong North Road, Minhang District". Both have the same ^8 (Xuhong North Road) and ^9 (Lane 18), but a conflict in ^2 (district). 2. Authoritative Verification: Query the "Shanghai Road Name-District Mapping Table" in A_data_sh to confirm that "the entire section of Xuhong North Road belongs to Xuhui District", and determine that the ^2 field (Minhang District) of record 2 is incorrect. 3. Forced Correction: Correct the ^2 field of record 2 to "Xuhui District" to unify the administrative affiliation of the two records, and proceed to the subsequent deduplication process. 4. Urban Renewal Adaptation: If the conflict originates from the old code, for example, record 1 is "Lane 1168, Yuanjing Road, Putuo District" (new code), and record 2 is "Lane 2138, Zhongshan North Road, Putuo District" (old code), query the "Shanghai Urban Renewal Address Update Table" in A_data_sh to adapt record 2... The old codes are linked to the new codes to unify administrative affiliation.

[0040] C. Shanghai Spatial Uniqueness Verification: Addressing the phenomenon of "same name, different address" in Shanghai (e.g., "Sunshine Building" exists in both Baoshan and Jinshan districts), the following verification is performed: 1. Multiple Attribution Confirmation: Querying A_data_sh confirms that "Sunshine Building" exists in two locations in Shanghai (Youyi Road Subdistrict, Baoshan District; Shihua Subdistrict, Jinshan District); 2. Shanghai High-Precision Geographic Decoding: Calling the high-precision geographic coding interface provided by the Shanghai Institute of Surveying and Mapping to obtain the latitude and longitude of the two "Sunshine Buildings" (e.g., Baoshan Sunshine Building: 121.48°E, 31.41°N; Jinshan Sunshine Building: 121.34°E, 30.74°N); 3. Distance Determination: The calculated distance between the two locations is approximately 85 kilometers, far exceeding the Shanghai localized preset threshold D_max (500 meters, adapted to the high-density address characteristics of Shanghai urban areas). Therefore, they are determined to be two independent entities, prohibiting merging and resolving the ambiguity of "same name, different address".

[0041] In this embodiment, the comprehensive fusion determination formula is: Where S_total is the total deduplication score calculated by the system, and W_vector is the similarity weight of the high-dimensional vector. Fusion is only performed when S_total exceeds the threshold T_fusion.

[0042] This invention also provides a high-precision deduplication and fusion database construction system based on hierarchical structured addresses, comprising the following modules: a structured module, used to perform preliminary word segmentation and preprocessing of address text using dictionary-based algorithm clustering and rule-based segmentation markers, and to perform sequence labeling using a deep learning model to obtain 18-level structured results; a conflict correction module, used to integrate authoritative data, maintain the authoritative data layer, and through conflict detection, authoritative verification, and forced correction, and old-to-new adaptation, associate the old code with the new code, unify administrative affiliation, and obtain Shanghai address records; and a dual anchoring module, used to define the relationship between urban and suburban anchor points, and to perform administrative-level pre-matching and verification. The anchor point matching module uses high-dimensional vector clustering for deduplication, BERT semantic similarity calculation, and a fusion determination formula to obtain the Shanghai address fusion result. The relationship mapping module integrates field information according to the principle of prioritizing official sources and supplementing non-NULL values ​​to form a complete address. It records the anchor point mapping relationship in the address database association table and stores aliases and old codes in specific fields to obtain the Shanghai address and anchor point mapping relationship. The continuous optimization module stores the processed address data in a dedicated database to build a complete Shanghai address database. It regularly updates the authoritative data layer and optimizes the model and algorithm according to new changes and business needs to obtain a continuously optimized address database.

[0043] This invention also provides a high-precision deduplication and database building device based on hierarchical structured addresses. This device may further include one or more power supplies, one or more wired or wireless network interfaces, one or more input / output interfaces, and / or one or more operating systems, such as Windows Server, MacOSX, Unix, Linux, FreeBSD, etc. Those skilled in the art will understand that the structure of the high-precision deduplication and database building device based on hierarchical structured addresses does not constitute a limitation on the computer device provided by this invention, and may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.

[0044] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. The computer-readable storage medium stores instructions that, when executed on a computer, cause the computer to perform the various steps of the high-precision deduplication and fusion database construction method based on hierarchical structured addresses provided in the above embodiments.

[0045] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

Claims

1. A high-precision deduplication and fusion database construction method based on hierarchical structured addresses, characterized in that, The high-precision deduplication and fusion database construction method based on hierarchical structured addresses includes the following steps: using a dictionary-based algorithm cluster and rule-based segmentation flag method to perform preliminary word segmentation and preprocessing on the address text, and using a deep learning model for sequence labeling to obtain 18-level structured results; By integrating authoritative data and maintaining an authoritative data layer, conflict detection, authoritative verification, and forced correction are performed. For old-to-new adaptation, old codes are associated with new codes to unify administrative affiliation, resulting in Shanghai address records. The anchor point relationship between urban and suburban areas is defined, and administrative-level pre-matching and core anchor point matching are performed. Through high-dimensional vector clustering for deduplication, BERT semantic similarity calculation, and fusion judgment formulas, the Shanghai address fusion result is obtained. Field information is integrated using the principle of prioritizing official sources and supplementing with non-NULL values ​​to form complete addresses. Anchor point mapping relationships are recorded in the address database association table, and aliases and old codes are stored in specific fields to obtain Shanghai addresses and anchor point mapping relationships. The processed address data is stored in a dedicated database to build a complete Shanghai address database. The authoritative data layer is updated regularly, and the model and algorithm are optimized based on new changes and business needs to obtain a continuously optimized address database.

2. The high-precision deduplication and fusion database construction method based on hierarchical structured addresses according to claim 1, characterized in that, The method employs a dictionary-based algorithm clustering and rule-based segmentation labeling to perform preliminary word segmentation and preprocessing on the address text, and uses a deep learning model for sequence labeling to obtain 18-level structured results. The steps include: using dictionary-based forward and reverse maximum matching, optimal matching, and rule-based segmentation labeling to perform preliminary word segmentation and preprocessing on the input unstructured Shanghai address text; using deep learning models such as Bi-LSTM-CRF and ID-CNN to accurately label the segmented sequence according to the 18-level labeling system and extract key information; and using domain-adaptive learning to transfer authoritative source resolution capabilities to low-quality data, integrating multiple models, and using a high-confidence priority and voting mechanism to obtain 18-level structured results for the Shanghai address.

3. The high-precision deduplication and fusion database construction method based on hierarchical structured addresses according to claim 1, characterized in that, The process of integrating authoritative data and maintaining the authoritative data layer involves conflict detection, authoritative verification, and forced correction. For urban renewal adaptation, the old code is associated with the new code to unify administrative affiliation and obtain Shanghai address records. This includes the following steps: integrating localized authoritative data in Shanghai using a combination of regular collection and proactive reporting. The authoritative data includes road names and districts, place names and districts, and urban renewal address update information. The authoritative data layer is updated and maintained in real time. By comparing low-level and high-level information of different address records, conflicts are detected. Verification is performed using accurate mapping relationships in the authoritative data layer, and erroneous fields are extracted. Erroneous fields are forcibly corrected to correct content to unify administrative affiliation. For addresses that have undergone administrative changes, the new codes associated with the old codes are extracted by querying the address update table, resulting in Shanghai address data with unified and accurate administrative affiliation.

4. The high-precision deduplication and fusion database construction method based on hierarchical structured addresses according to claim 1, characterized in that, The process involves defining the anchor point relationship between urban and suburban areas, performing administrative-level pre-matching and core anchor point matching, and obtaining the Shanghai address fusion result through high-dimensional vector clustering for deduplication, BERT semantic similarity calculation, and a fusion judgment formula. This includes the following steps: defining anchor point relationships in Shanghai's urban areas using lane numbers as official codes and courtyards as colloquial terms, and in suburban areas using village numbers as official codes and natural villages as colloquial terms; employing administrative-level pre-matching, prioritizing the verification of the first four levels of address fields and correcting them to standard administrative affiliation; performing core anchor point matching, extracting address entity features through high-dimensional vector clustering for deduplication, and calculating the semantic similarity of core anchor point fields using a BERT twin network model; combining the fusion judgment formula to obtain the total deduplication score, and then obtaining the address's latitude and longitude through geographic decoding, calculating the distance, and determining those meeting the distance condition as physically overlapping addresses, thus obtaining the accurately fused Shanghai address result.

5. The high-precision deduplication and fusion database construction method based on hierarchical structured addresses according to claim 1, characterized in that, The process of integrating field information using the principle of prioritizing official sources and supplementing with non-NULL values ​​to form a complete address, recording anchor mapping relationships in the address database association table, and storing aliases and old codes in specific fields, yields Shanghai addresses and anchor mapping relationships. This includes the following steps: First, using the principle of prioritizing official Shanghai sources and supplementing with non-NULL values, field information from different data sources is integrated, prioritizing fields from official Shanghai data sources. Second, when official data is empty, supplementary data is extracted from non-empty data sources to form a complete address. Third, by analyzing the official codes and colloquial names in the address, the anchor mapping relationships are recorded in detail in the address database association table. Fourth, unique aliases and old codes of Shanghai addresses are extracted and stored in specific fields, thereby improving the address recognition rate in local scenarios.

6. The high-precision deduplication and fusion database construction method based on hierarchical structured addresses according to claim 1, characterized in that, The process of storing processed address data in a dedicated database to build a complete Shanghai address database, regularly updating the authoritative data layer, and optimizing models and algorithms based on new changes and business needs to obtain a continuously optimized address database includes the following steps: Using a database management module, the Shanghai address data, after word segmentation, structuring, conflict correction, and fusion, is accurately stored in a specially constructed database to build a complete Shanghai address database; By setting scheduled tasks and a real-time monitoring mechanism, the latest data is regularly obtained from authoritative channels to update the authoritative data layer; Based on newly emerging address changes and constantly evolving business needs, targeted optimizations are performed, continuously adjusting parameters and strategies to obtain Shanghai address deduplication and fusion results with continuously improving accuracy and effectiveness.

7. The high-precision deduplication and fusion database construction method based on hierarchical structured addresses according to claim 4, characterized in that, The formula for determining fusion is as follows: Where S_total is the total deduplication score calculated by the system, and W_vector is the similarity weight of the high-dimensional vector. Fusion is only performed when S_total exceeds the threshold T_fusion.

8. A high-precision deduplication and fusion database construction system based on hierarchical structured addresses, characterized in that: The high-precision deduplication and fusion database building system based on hierarchical structured addresses includes the following modules: a structured module, which uses a dictionary-based algorithm cluster and rule-based segmentation flag method to perform preliminary word segmentation and preprocessing on the address text, and uses a deep learning model for sequence labeling to obtain 18-level structured results; The conflict correction module is used to integrate authoritative data and maintain the authoritative data layer. Through conflict detection, authoritative verification and forced correction, the old code is associated with the new code for old-to-new adaptation, and the administrative affiliation is unified to obtain the Shanghai address record. The dual anchoring module is used to define the relationship between urban and suburban anchor points, perform administrative layer pre-matching and core anchor point matching, and obtain the Shanghai address fusion result through high-dimensional vector clustering deduplication, BERT semantic similarity calculation and fusion judgment formula. The relation mapping module integrates field information using the principle of prioritizing official sources and supplementing with non-NULL values ​​to form a complete address. It records anchor mapping relationships in the address database association table, stores aliases and old codes in specific fields, and obtains Shanghai addresses and anchor mapping relationships. The continuous optimization module stores the processed address data in a dedicated database, builds a complete Shanghai address database, regularly updates the authoritative data layer, and optimizes the model and algorithm based on new changes and business needs to obtain a continuously optimized address database.

9. A high-precision deduplication and fusion database building device based on hierarchical structured addresses, characterized in that: The high-precision deduplication and fusion database building device based on hierarchical structured addresses includes a memory and at least one processor. The memory stores instructions, and the at least one processor calls the instructions in the memory to cause the high-precision deduplication and fusion database building device based on hierarchical structured addresses to perform each step of the high-precision deduplication and fusion database building method based on hierarchical structured addresses as described in any one of claims 1-7.

10. A computer-readable storage medium storing instructions thereon, characterized in that, When the instructions are executed by the processor, they implement each step of the high-precision deduplication and fusion library construction method based on hierarchical structured addresses as described in any one of claims 1-7.