Hybrid chinese address resolution method based on sequence labeling and double index

CN122735686APending Publication Date: 2026-09-11HENAN CRAFTSMANSHIP INFORMATION TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610961736.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-30
Publication Date
2026-09-11

AI Technical Summary

Technical Problem

[0002]中文地址是电子商务物流、快递配送、位置服务、政务人口管理、金融反欺诈及客户关系管理等业务中的基础数据,上述业务通常需要将用户输入的非结构化中文地址文本转换为包含省、市、区县、街道、道路、门牌号等字段的结构化层级数据,以便进行配送分单、区域定位、数据归档及风险校验,然而,中文地址书写方式具有较强的不规范性和口语化特征,常见情形包括省市后缀省略、别名或简称使用、同名行政区跨区域存在、直辖市层级跳转以及地级市直管镇等特殊行政区划结构,导致地址文本的准确解析和标准化归一化难度较高

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122735686A_ABST
    Figure CN122735686A_ABST
Patent Text Reader

Abstract

The application discloses a hybrid Chinese address resolution method based on sequence labeling and double-layer index, and comprises the following steps: S1, loading standard administrative division data, and constructing a double-layer memory index, wherein the double-layer memory index comprises a global inverted hash index for mapping candidate administrative area nodes by using keywords, and a hierarchical Trie tree for organizing nodes according to administrative parent-child levels and pre-calculating ancestor bitmaps; S2, receiving a Chinese address text to be resolved, and recalling keyword matching items and candidate administrative area nodes thereof through an AC automatic machine on a cleaning view for retaining original text offset mapping. The hybrid Chinese address resolution method based on sequence labeling and double-layer index can better adapt to complex scenes such as non-standard Chinese addresses, administrative areas with the same name and special administrative level structures by combining a double-layer memory index, an AC automatic machine recall, ancestor bitmap blood relationship determination, strong anchor point trimming, path score completion and a sequence labeling model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of Chinese address processing technology, specifically a hybrid Chinese address parsing method based on sequence labeling and two-level indexing. Background Technology

[0002] Chinese addresses are fundamental data in e-commerce logistics, express delivery, location services, government population management, financial anti-fraud, and customer relationship management. These businesses typically need to convert unstructured Chinese address text input by users into structured hierarchical data containing fields such as province, city, district / county, street, road, and house number for delivery order allocation, regional positioning, data archiving, and risk verification. However, the writing style of Chinese addresses is highly irregular and colloquial. Common issues include the omission of province / city suffixes, the use of aliases or abbreviations, the existence of administrative regions with the same name across regions, jumps between municipalities, and special administrative division structures such as towns directly under prefecture-level cities. This makes accurate parsing and standardization of address text quite difficult.

[0003] Existing Chinese address parsing technologies mainly include rule-based or regular expression-based parsing methods, place name dictionary-based or string matching parsing methods, deep learning sequence labeling parsing methods, and methods that rely on administrative division databases or place name dictionaries for recall followed by rule disambiguation. Among these, rule-based or dictionary matching methods are simple to implement and have a relatively fast processing speed, but they usually rely on manual rules or dictionary coverage and have poor adaptability to non-standard address expressions such as abbreviations, aliases, no suffixes, contractions, and inversions. Deep learning sequence labeling methods can identify entity fragments in address text and have a certain semantic generalization ability, but their output results are usually still mainly text fragments, making it difficult to directly bind to standard administrative division codes. Furthermore, they still have limitations in terms of processing cost and response latency in special administrative hierarchical structures and high-concurrency online parsing scenarios.

[0004] In practical applications, existing methods still struggle to simultaneously achieve parsing accuracy, normalization capability, and online service performance. For example, numerous administrative regions with the same name exist nationwide, making simple dictionary matching prone to generating multiple candidate nodes and failing to determine a unique standard region. Even if sequence labeling models can identify entities such as "Zhengzhou" and "Jinshui," they typically cannot directly provide the corresponding standard administrative division code. Some schemes lack explicit modeling of administrative hierarchy parent-child and ancestor relationships, making it difficult to use identified higher-level administrative regions to constrain lower-level region matching. Furthermore, special administrative division structures such as municipalities, prefecture-level city-administered towns, and autonomous prefectures can easily lead to discontinuous parsing paths or normalization failures. In addition, traditional schemes still have shortcomings in reliability scoring, online updates of administrative division data, and high-concurrency, low-latency services. Therefore, it is necessary to provide a Chinese address parsing scheme that combines candidate recall, hierarchical relationship constraints, path scoring, and fine-grained entity recognition. Summary of the Invention

[0005] The purpose of this invention is to provide a hybrid Chinese address parsing method based on sequence labeling and two-level indexing to solve the problems mentioned in the background art.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a hybrid Chinese address resolution method based on sequence labeling and two-level indexing, executed by the processor of the address resolution service, comprising: S1. Load standard administrative division data and construct a two-layer memory index, which includes a global inverted hash index that maps candidate administrative division nodes with keywords, and a hierarchical Trie tree that organizes nodes according to administrative parent-child levels and pre-computes the ancestor bitmap. S2. Receive the Chinese address text to be parsed, and recall keyword matching items and their candidate administrative region nodes on the cleaned view that retains the original offset mapping through the AC automaton; S3. Use the ancestor bitmap to determine the same, ancestor or descendant relationship between candidate nodes, and suppress the included matching items only when the text range contains and the relationship exists; S4. Based on candidate uniqueness, candidate node level and neighboring ancestor matching, determine strong anchor points, and prune multiple candidate nodes according to strong anchor points and ancestor bitmaps, and construct candidate address paths along the same lineage path; S5. The candidate address paths are comprehensively scored based on hierarchical completeness, hierarchical continuity, strong anchor point ratio, physical continuity, text coverage, and reverse order penalty. The address paths that meet the threshold are selected and the missing administrative levels are filled in. S6. Use a sequence labeling model to identify fine-grained address entities outside administrative regions, merge the normalized administrative region entities generated by the address path with the fine-grained address entities according to their original positions, and output the structured address parsing results.

[0007] As a further step, the administrative division node includes a region code, standard name, abbreviation, alias, level, and parent node code; the standard name, abbreviation, and alias are all written as keywords into the global inverted hash index, and the ancestor bitmap is generated by superimposing the parent node's ancestor bitmap with the parent node's identifier.

[0008] As a further step, the cleaned view filters out numbers, punctuation marks, and special characters while retaining Chinese and English letters, while the original text view retains the complete text; the matching position of the AC automaton on the cleaned view is backfilled to the original text position through offset mapping.

[0009] As a further step, the suppression of included matches includes: sorting the matches in ascending order of their start position and descending order of their end position; when a short match is included by a long match, and the candidate nodes of the two are the same or in the same lineage path, deleting the short match; otherwise, retaining it.

[0010] As a further step, the pruning of multiple candidate nodes includes: collecting candidate nodes and their ancestors to form a set of context nodes, and using the unique candidate as a strong anchor point; prioritizing the retention of candidates with strong anchor points as ancestors or as ancestors of strong anchor points, and continuing convergence according to the deepest strong anchor point if necessary.

[0011] As a further step, the pruning of multiple candidate nodes also includes fragment deduplication, position pruning, and prefix chain pruning; fragment deduplication uses rolling hashing, position pruning selects the best candidates based on lineage support, keyword length, candidate uniqueness, and hierarchical information content, and prefix chain pruning removes candidates that have no lineage relationship with the confirmed upper-level nodes.

[0012] As a further step, the construction of candidate address paths includes: determining strong anchor point scores based on candidate level not lower than the city level, candidate uniqueness, and nearby ancestor matching; using the strongest anchor point at the highest level as the main anchor point, collecting its ancestor matching items and descendant matching items, and sorting them by administrative level to form candidate address paths.

[0013] As a further step, the hierarchical continuity is considered as continuous across levels from the provincial level to the district / county level in the municipality mode, and as continuous across levels from the city level to the township / street level in the prefecture-level city directly administered town mode; the completion node is determined by the ancestor bitmap of the deepest node and marked as (-1,-1), and does not participate in the calculation of text coverage, physical continuity and reverse order penalty.

[0014] As a further step, the sequence labeling model includes a Chinese pre-trained encoder, a bidirectional long short-term memory network, a linear mapping layer, and a conditional random field layer; it identifies entities such as roads, house numbers, residential areas or communities, buildings, units, floors, room numbers, points of interest, sub-points of interest, details, and village / group entities, and skips the administrative region entities it outputs during merging and prioritizes retaining normalized administrative region entities.

[0015] Furthermore, it also includes non-stop hot updates: based on the new standard administrative division data, a new in-memory data connection, a two-level in-memory index, an AC automaton, a coarse-ranking matcher, and a fine-ranking decision maker are built; if the build is successful, the runtime references are replaced once, and if the build fails, the old version is maintained and the service continues.

[0016] Compared with existing technologies, the beneficial effects of this invention are: by combining dual-layer memory indexing, AC automaton recall, ancestor bitmap lineage determination, strong anchor point pruning, path scoring completion, and sequence labeling model, the Chinese address parsing process has the ability to normalize administrative divisions, identify fine-grained addresses, and provide online services. It can better adapt to complex scenarios such as non-standard Chinese addresses, administrative regions with the same name, and special administrative hierarchical structures, as detailed below.

[0017] 1. Improve the accuracy and normalization stability of administrative region resolution by quickly determining the same, ancestral, or descendant relationships between candidate nodes using a hierarchical Trie tree and ancestor bitmap. Based on this, suppress only matches that satisfy interval inclusion and have a lineage relationship. Combine strong anchor points, lineage pruning, path scoring, and path completion to select the final address path. This can reduce mismatches in cases of districts and counties with the same name, omitted abbreviations, and missing suffixes, enabling administrative region entities to be directly bound to standard area codes and improving the usability of structured address results.

[0018] 2. Balancing the standardization of administrative regions with the semantic recognition of detailed addresses, the system uses index retrieval and path ranking to determine administrative region entities, while using sequence labeling models to identify non-administrative region fine-grained entities such as roads, house numbers, residential areas, buildings, units, room numbers, and points of interest. During the merging phase, administrative region entities output by the sequence labeling model are skipped, and normalized administrative region entities are retained first. This avoids the problem that pure sequence labeling models can only output text fragments and are difficult to directly normalize to standard administrative regions, while also reducing duplication and conflicts between administrative region entities and detailed address entities.

[0019] 3. Improve online parsing performance and data update availability by loading administrative division data, inverted index, hierarchical Trie tree, ancestor bitmap, and AC automaton into memory. During parsing, there is no need to rely on external databases for sequential queries, and the runtime lineage determination cost is reduced by pre-compiling the ancestor bitmap. At the same time, when updating administrative division data, a new state is built first, and then the runtime reference is replaced as a whole. When the new index or matcher fails to build, the old state is retained to continue to provide services. This helps to reduce online parsing latency and improve the continuity of services in the scenario of administrative division adjustment. Attached Figure Description

[0020] Figure 1 This is a schematic diagram of the system architecture of the present invention; Figure 2 This is a flowchart of the coarse matching process of the present invention; Figure 3 This is a schematic diagram illustrating the feasibility of the cutting process of this invention; Figure 4 This is a schematic diagram of the candidate address path construction process of the present invention; Figure 5 This is a schematic diagram of the path scoring and threshold filtering process of the present invention; Figure 6 This is a schematic diagram of the path completion process of the present invention. Detailed Implementation

[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0022] Please see Figures 1-6 The present invention provides the following technical solution: The hybrid Chinese address parsing method based on sequence labeling and two-level indexing includes the following steps: S1. Load standard administrative division data and construct a two-level in-memory index. The hybrid Chinese address resolution method based on sequence labeling and two-level indexing can be executed by the address resolution service's processor. The address resolution service can be deployed on a single server, server cluster, cloud server, or cloud server cluster. It is used to segment unstructured free text addresses into structured fields such as province, city, district / county, township / street, road, and house number, and normalize the administrative division entities within them to standard administrative division nodes.

[0023] In this embodiment, step S1 of the method specifically involves: loading standard administrative division data and constructing a two-level memory index. The two-level memory index includes a global inverted hash index that maps candidate administrative division nodes by keywords, and a hierarchical Trie tree that organizes nodes according to administrative parent-child levels and pre-computes the ancestor bitmap.

[0024] Specifically, in this embodiment, standard administrative division data is read when the address resolution service starts or during hot data updates. This data can be stored in object storage as a table and managed using a versioned table format. In this embodiment, ACID transaction capabilities and versioning are provided by DeltaTable, Iceberg, Hudi, or other versioned table formats. Object storage serves only as the underlying storage medium; ACID capabilities are not a necessary limitation. After reading the remote table dataset, the address resolution service registers it in an in-memory database and establishes a secondary index for fast querying. All administrative division data is loaded into the process memory, eliminating the need for subsequent resolution processes to access external databases online or perform disk I / O on the resolution request path.

[0025] Specifically, each administrative division node in the standard administrative division data is encapsulated as an AreaNode. An AreaNode is used to represent a single administrative division node and includes at least the area code, standard name, main abbreviation, administrative level, parent node code, alias array, child node set, path code, and ancestor bitmap. In this embodiment, name represents the main abbreviation and floor represents the administrative level.

[0026] Optionally, the data fields of AreaNode can be organized according to the following table: Specifically, the global inverted hash index is GlobalHashIndex, which maintains a keyword→List. <areanode>is a many-to-one mapping, and keywords include keywords in the title, name and alias_names of administrative division nodes. For example, the title may be Henan Province, the name may be Henan, and alias_names may include aliases such as Yu. A single keyword can correspond to one or more administrative division nodes; when a single keyword corresponds to multiple administrative division nodes, the keyword forms multi-candidate recall results, which are disambiguated through subsequent steps of血缘 suppression, strong anchor clipping, path scoring and path completion.

[0027] Further, entries in the alias array shall be limited to canonical aliases, abbreviations, historical names or business-maintained place name mappings that can point to corresponding administrative division nodes; if a non-administrative division POI name such as Zhengzhou Railway Station is used as a business mapping word, its mapping source and the pointed administrative division node shall be recorded in the data, and it shall be distinguished from POI entities in the subsequent entity merging stage, so as to avoid equating ordinary point of interest names directly with administrative division names. Through this processing method, the interference of non-administrative division names on the administrative division normalization results can be reduced while retaining the business alias recall capability.

[0028] Specifically, the hierarchical Trie tree is HierarchicalTrie, which takes parent_id as parent-child pointers and organizes all AreaNodes into a multi-way tree including provincial, municipal, district / county and township / street levels. In this hierarchical Trie tree, provincial nodes are higher-level nodes, municipal nodes are hung under corresponding provincial nodes, district / county-level nodes are hung under corresponding municipal nodes, and township / street-level nodes are hung under corresponding district / county-level nodes or special directly-administered nodes. For special administrative division structures such as municipalities directly under the Central Government and towns directly administered by prefecture-level cities, this embodiment does not require that each level must actually appear in the input text, but performs compatible processing through subsequent path scoring and path complement mechanisms.

[0029] Further, after the loading of the hierarchical Trie tree is completed, the address resolution service calls build_path_codes() to pre-calculate path_code and ancestor_bitmap for each administrative division node. Wherein, path_code is a slash path encoding in the form of / 410000 / 410100 / 410105 / ; ancestor_bitmap is a set or bitmap representation of all ancestor nodes of the current node, which is used to quickly determine whether any two administrative division nodes have an ancestor, descendant or same-node relationship.

[0030] Specifically, let the administrative division node be x, its unique code is x.area_id, and its ancestor set / bitmap is B(x), then: B(x) = {y.area_id | y ∈ path(x), y.area_id ≠ x.area_id}, where path(x) represents the complete path from the root node to node x. This expression defines the generation boundary of ancestor_bitmap, that is, the ancestor bitmap records the parent nodes of node x, but not node x itself.

[0031] For any two nodes a and b, the rules for determining blood relations are as follows: isAncestor(a,b)=1⇔a.area_id∈B(b); isDescendant(a,b)=1⇔b.area_id∈B(a); `sameLineage(a,b) = (a.area_id = b.area_id) OR isAncestor(a,b) OR isAncestor(b,a)`; where `isAncestor(a,b)` indicates whether node `a` is an ancestor of node `b`, `isDescendant(a,b)` indicates whether node `a` is a descendant of node `b`, and `sameLineage(a,b)` indicates whether node `a` and node `b` are on the same administrative region lineage path. The same administrative region lineage path includes three cases: they are the same node, node `a` is an ancestor of node `b`, and node `b` is an ancestor of node `a`.

[0032] In engineering implementation, ancestor_bitmap can be set. <int>Storage can be achieved using methods such as BitSet, uint64 array, or RoaringBitmap. If BitSet encoding is used, all area_ids must first be mapped to consecutive bit indices idx(area_id), with parent-child nodes satisfying the following: The code `Bitmap(child) = Bitmap(parent) OR bit(idx(parent))` determines whether `a` is an ancestor of `b` at runtime. This only requires checking if the bit at the `idx(a)` position in `Bitmap(b)` is 1. If a set implementation is used, then the member check of `a.area_id∈B(b)` is performed. This approach avoids backtracking along the parent node pointer at runtime, thus supporting subsequent lineage-based NMS, candidate pruning, path scoring, and path completion.

[0033] Furthermore, ancestor_bitmap is not just an ordinary auxiliary field, but a core lineage determination structure that runs through coarse-sorting suppression, fine-sorting disambiguation, and path completion. For nested, abbreviated, or homonymous address fragments such as Beijing / Beijing Municipality, Henan / Henan Province, and Chaoyang / Chaoyang District, the address resolution service can determine whether different candidate administrative region nodes have the same node, ancestor, or descendant relationship based on ancestor_bitmap, without needing to re-query the database or traverse the complete administrative division tree in each resolution request.

[0034] Furthermore, the address resolution service can also construct an AC automaton on the keyword set of GlobalHashIndex, which is used for simultaneous scanning of multiple pattern strings in the Chinese address text to be resolved during the subsequent coarse-ranking matching stage. In this embodiment, the AC automaton is not a component level of the two-level memory index; the two-level memory index body consists of GlobalHashIndex and HierarchicalTrie, and the AC automaton is a coarse-ranking matching structure generated based on the keyword set of GlobalHashIndex, which together with the two-level memory index forms the runtime address resolution index system.

[0035] Specifically, the GlobalHashIndex, HierarchicalTrie, and AC automaton work together at runtime: the AC automaton is responsible for simultaneously locating all possible keyword positions in the text; the GlobalHashIndex is responsible for determining which candidate administrative region nodes correspond to a keyword; and the HierarchicalTrie and its ancestor_bitmap are responsible for determining the parent-child, ancestor, or descendant relationships between these candidate administrative region nodes. In other words, the AC automaton solves the problem of where the keyword is hit, the GlobalHashIndex solves the problem of which administrative region nodes the keyword might correspond to, and the HierarchicalTrie and ancestor_bitmap solve the problem of whether these candidate nodes belong to the same administrative region lineage.

[0036] Furthermore, after S1 is completed, a data structure is formed in memory that includes at least the following runtime states: a set of standard administrative division nodes, a global inverted hash index from keywords to a list of candidate administrative division nodes, a hierarchical Trie tree organized by parent-child administrative levels, a path_code for each node, an ancestor_bitmap for each node, and an AC automaton constructed from the keyword set. These runtime states can be mounted in the global state object of the address resolution service for subsequent S2 to S6 resolution processes to call.

[0037] In summary, S1 pre-solidifies keyword recall and administrative hierarchy lineage constraint capabilities by loading standard administrative division data into memory and constructing two complementary indexes: GlobalHashIndex and HierarchicalTrie. GlobalHashIndex ensures that abbreviations, full names, and aliases can be quickly recalled, while HierarchicalTrie and ancestor_bitmap ensure that the recalled results can be quickly determined to be the same node, ancestor, descendant, or unrelated. Therefore, in subsequent processing scenarios such as administrative regions with the same name, abbreviation omission, municipalities crossing levels, and towns directly under prefecture-level cities, this embodiment no longer relies solely on pure rule or pure sequence labeling models. Instead, it can utilize the parent-child lineage relationship in the standard administrative division system to constrain candidate results, providing a unified data foundation for subsequent lineage NMS, strong anchor pruning, candidate address path construction, path scoring, and path completion.

[0038] S2. Receive the Chinese address text to be parsed, perform dual-view preprocessing, AC recall, and lineage NMS suppression. Specifically, step S2 is as follows: Receive the Chinese address text to be parsed, recall keyword matching items and their candidate administrative region nodes on the cleaned view that retains the original text offset mapping through the AC automaton; use the ancestor bitmap to determine the same, ancestor, or descendant relationship between candidate nodes, and suppress the included matching items only when the text range contains the matching items and a relationship exists.

[0039] Specifically, after receiving the Chinese address text to be parsed, the address resolution service does not directly and irreversibly delete the original text. Instead, it generates two parallel views. The first view is a coarse-grained administrative division view, filtering out numbers, punctuation, spaces, and special characters, retaining only Chinese characters and English letters, for coarse-grained matching of administrative division keywords. The second view is the original text view, retaining the complete input text, for subsequent fine-grained entity recognition such as roads, house numbers, buildings, units, room numbers, and POIs. In this embodiment, the cleaned view is only used for coarse-grained recall of administrative divisions and is not used as the final entity extraction text, nor is the number, house number, and building / room number information in the original text view deleted.

[0040] Furthermore, when generating the coarse-grained administrative division view, the address resolution service simultaneously establishes an offset mapping between the cleaned character positions and the original character positions, enabling the start / end offsets obtained subsequently through the AC automaton to be backfilled into the original text view. This embodiment uniformly adopts a left-closed, right-open interval [start, end), where start represents the starting character position of the hit fragment in the original text, and end represents the position after the ending character of the hit fragment in the original text. If the engineering implementation uses closed intervals, consistency should be maintained throughout the recall, suppression, scoring, and merging processes to avoid mixing interval definitions in the same process.

[0041] Specifically, the AC automaton is constructed based on the keyword set of GlobalHashIndex in S1. Keywords include the standard name, primary abbreviation, and alias of administrative division nodes. The AC automaton scans all keyword-matched items in the coarse-ranked administrative division view at once. For each match, it retrieves a list of candidate administrative division nodes from GlobalHashIndex based on the keyword, thus forming a coarse-ranked match. The AC automaton solves the problem of where the keyword is matched, GlobalHashIndex solves the problem of which candidate administrative division nodes correspond to the keyword, and HierarchicalTrie and ancestor_bitmap solve the problem of whether there are lineage relationships between candidate administrative division nodes.

[0042] The table below shows the coarse-sorted matching fields: The coarse-ranked matching field table described above is used to illustrate the data boundaries of the AC automaton's recall results. In other words, the AC automaton itself is only responsible for multi-pattern string scanning; the candidate nodes, region codes, parent-child relationships, and ancestor bitmaps required for administrative region normalization are still provided by the two-level in-memory index constructed by S1.

[0043] Furthermore, the address resolution service performs lineage-based stack-based nonmaximum suppression on all matches output by the AC automaton. For all matches, they are first sorted in ascending order of start and descending order of end, while a stack structure is used to maintain the current long matches that may contain subsequent short matches. Let the short match be m_i=(s_i,e_i,C_i), and the long match be m_j=(s_j,e_j,C_j), where s and e represent the start and end positions of the match in the original text, respectively, and C represents the set of candidate administrative region nodes recalled by the keyword. Short match m_i is suppressed only when both of the following conditions are met: The interval contains the following conditions: s_j≤s_i and e_i≤e_j; Bloodline-related conditions: NodeIds(m_i)∩NodeIds(m_j)≠∅ or NodeIds(m_i)∩Ancestors(m_j)≠∅ or NodeIds(m_j)∩Ancestors(m_i)≠∅.

[0044] Where NodeIds(m) represents the set of candidate node IDs for match m, and Ancestors(m) represents the set of ancestor IDs for all candidate nodes of match m. The above lineage-related conditions respectively indicate that two matches point to the same administrative region node, the shorter match is an ancestor of a candidate node of the longer match, or the longer match is an ancestor of a candidate node of the shorter match.

[0045] The table below shows the bloodline NMS inhibition conditions: The aforementioned lineage-based NMS suppression condition table is used to illustrate that the NMS in this embodiment is not a traditional pure interval NMS. Traditional interval NMS deletes entries as long as the short interval is contained within the long interval, which can easily lead to the accidental deletion of administrative region keywords embedded in road names, school names, park names, and point-of-interest names. In this embodiment, when determining whether a short match should be suppressed by a long match, in addition to requiring that the short matching interval be contained within the long matching interval, it also requires that the candidate administrative region nodes recalled by both have the same node, ancestor, or descendant relationship.

[0046] For example, regarding Chaoyang District of Beijing, the short match "Beijing" and the long match "Beijing" not only have a range inclusion relationship, but their candidate nodes also belong to the same administrative district lineage. Therefore, "Beijing" can be suppressed, preserving a more complete picture of Beijing. However, for "Beijing University of Science and Technology Road" or "Beijing University of Science and Technology Road in Beijing," where "Beijing University of Science and Technology Road" is more likely a road name or POI name, if the long segment does not point to an administrative district node, or if its candidate nodes do not have a lineage relationship with the candidate nodes of "Beijing," then even if the text range has an inclusion relationship, "Beijing" will not be suppressed. Therefore, the incorrect removal of nested patterns without lineage can be avoided.

[0047] In summary, S2 combines dual-view preprocessing, AC automaton recall, and lineage NMS suppression to ensure the speed of recalling administrative division keywords while preserving the numerical and positional offset information in the original text that is crucial for fine-grained address parsing. Furthermore, by adding administrative division lineage constraints to the text inclusion relationship through ancestor bitmaps, it enables the elimination of irrelevant nested segments during the coarse-sorting stage, providing a cleaner candidate matching set for subsequent feasibility pruning and path construction.

[0048] S3. Perform feasibility pruning based on strong anchor points and ancestor bitmaps, and construct candidate address paths. Specifically, step S3 of the method is as follows: determine strong anchor points based on candidate uniqueness, candidate node level and neighboring ancestor matching, prune multiple candidate nodes based on strong anchor points and ancestor bitmaps, and construct candidate address paths along the same lineage path.

[0049] Specifically, after obtaining the Match list output by S2, the address resolution service first enters the feasibility pruning stage. The purpose of feasibility pruning is not to directly select the final path, but to use the identified strong context information and administrative division relationships to preemptively remove candidate nodes that are obviously unlikely to belong to the current address path, thereby reducing the combinatorial complexity of subsequent path construction and scoring.

[0050] Furthermore, the address resolution service traverses the coarse-ranked Match list, collects the area_id of all candidate nodes and their ancestor sets, and constructs the context node set Context: Context=⋃{m∈M}⋃{c∈C(m)}({c.area_id}∪B(c)), where M is the coarse-ranked matching set, C(m) is the set of candidate administrative region nodes for matching item m, and B(c) is the ancestor set of node c, i.e., the ancestor_bitmap pre-computed in S1.

[0051] Simultaneously, matches that only recall a single candidate are considered strong context anchors at that matching level, forming a set of strong anchors A, and their encoding set Id(A) and ancestor set Anc(A) are recorded: A={c|c∈C(m),m∈M,|C(m)|=1}; Id(A) = {a.area_id | a∈A}; Anc(A) = ⋃_{a∈A}B(a).

[0052] In this embodiment, the set of strong anchor points A is not limited to retaining only one node at each level; when there are multiple unique candidate strong anchor points that are not related to each other in the same text, they can be registered separately according to text fragments or prefix chains to avoid mistakenly removing subsequent addresses in scenarios where multiple addresses appear consecutively.

[0053] Specifically, the address resolution service first performs fragment deduplication. For periodically repeating fragments such as "XX province XX city XX province XX city" caused by user error, a rolling hash function can be used to detect duplicate windows and delete redundant matches corresponding to the duplicate fragments. Fragment deduplication is used to reduce path duplication caused by repeated inputs without changing the candidate node set of non-duplicate fragments.

[0054] Furthermore, lineage pruning is performed on multiple candidate matches. Lineage pruning is only applied to matches where |C(m)|>1; for unique candidate matches, which already have high determinism, they are usually retained directly. For any candidate node c in a multiple candidate match, the retention condition is: keep(c)=(c.area_id∈Id(A))OR(B(c)∩Id(A)≠∅)OR(c.area_id∈Anc(A)), where c.area_id∈Id(A) indicates that the candidate node and the strong anchor point are the same node; B(c)∩Id(A)≠∅ indicates that the candidate node has a strong anchor point as its ancestor; and c.area_id∈Anc(A) indicates that the candidate node itself is an ancestor of a strong anchor point. Compared to simply retaining nodes that are ancestors of strong anchors or are ancestors of strong anchors, this embodiment adds a retention condition where the candidate node and the strong anchor are the same node, in order to avoid nodes from the same administrative region being mistakenly pruned because they are not considered their own ancestors.

[0055] If multiple candidates still exist after retaining them according to the above conditions, the deepest strong anchor point a is selected for further convergence, and only candidates that satisfy a.area_id∈B(c) are retained. If no candidate satisfies the retention condition, the process reverts to the conservative rule that the ancestor appears in the context: B(c)∩(Context−{c.area_id})≠∅. If the process is still empty after reverting, all original candidates of the matching item are retained to prevent excessive pruning from prematurely deleting the correct candidate.

[0056] The following table shows the feasibility tailoring process: The above feasibility pruning flowchart is used to illustrate that the fine-ranking preprocessing in this embodiment is not a single rule filter, but rather combines context, repeated fragments, administrative lineage, text position and prefix chain consistency to perform candidate pruning, thereby reducing the interference of administrative regions with the same name and irrelevant place name fragments on subsequent path construction.

[0057] Furthermore, the address resolution service performs position pruning on candidate matches where the text interval overlap length is greater than one character. Position pruning does not simply retain the longest string, but comprehensively considers lineage support, keyword length, candidate uniqueness, and administrative level information. For any match m, the information content score is defined as: InfoScore(m) = LineageScore(m) + LengthScore(m) + UniquenessScore(m) + LevelScore(m). The table below shows the breakdown of information score items: The above information content score table is used to explain the criteria for selecting the best location for cropping.

[0058] Furthermore, during location pruning, the address resolution service traverses overlapping matches in ascending order of start and descending order of end, retaining those with higher InfoScores and marking the lower ones as removed. When scores are equal, the preceding match is retained, i.e., the one with an earlier start and a longer interval. Therefore, it avoids retaining erroneous segments based solely on string length and also avoids deleting high-confidence administrative region segments due to an excessive number of candidates.

[0059] Furthermore, the address resolution service performs prefix chain pruning. Prefix chain pruning maintains the confirmed administrative region prefix chains from left to right according to the text. The prefix chain can include the confirmed prefix code set `confirmed_prefix_ids`, the confirmed nodes at each level `nodes_by_floor`, the pending nodes at each level `pending_by_floor`, the confirmed chain length `prefix_chain_length`, and the historical maximum level `prev_max_floor`. When the chain length reaches 2, if a new candidate has no lineage relationship with its upper-level confirmed nodes or confirmed prefix set, it is first placed into the pending set; if a confirmed node appears at the same level afterward, the pending candidate is pruned.

[0060] Furthermore, prefix chain pruning also supports legal path jumps. When the following conditions are met: current_floor=1 AND prev_max_floor≥2, that is, after the text has already identified lower-level levels such as city, district, and street, and a new provincial node appears, the address resolution service determines that the text may enter a second independent address or a new administrative path. At this time, the system clears the current prefix chain, the pending set, and the confirmed sets of each level, and restarts chain construction from the provincial node. Therefore, prefix chain pruning can both remove noisy candidates unrelated to the confirmed path within a single address and not block legal provincial jumps when multiple addresses appear consecutively in the same text.

[0061] After completing the feasibility trimming, the address resolution service enters the path construction phase. The path construction phase first calculates a strong anchor score for each match m. Let c0 be the preferred candidate node for this match, C(m) be its candidate set, and define the nearest ancestor indicator NearAncestor(m): when there exists another match m', whose candidate node's area_id belongs to the ancestor set of c0, and the original distance between them |m.start−m'.end|≤20 characters, NearAncestor(m)=1; otherwise, NearAncestor(m)=0. The strong anchor score is calculated as: AnchorScore(m)=I(c0.floor≥2)+I(|C(m)|=1)+2×I(NearAncestor(m)=1), where I(condition) is the indicator function, taking 1 when the condition is true and 0 otherwise. When AnchorScore(m)≥3, the match m is marked as a strong anchor. This scoring system reflects three types of high-confidence signals: the candidate node level is no lower than the city level, the keyword recall result is unique, and there is an ancestor nearby that can form the same lineage. The nearest ancestor matching weight is 2, so that the combination of lower-level node + nearby upper-level context can reach the threshold even if the candidate is not unique, and thus be identified as a strong anchor point.

[0062] In this embodiment, c0 should be the preferred candidate node after feasible pruning; if a matching item still retains multiple candidate nodes after pruning, the anchor point score corresponding to each candidate node can be calculated separately, and the candidate node with the highest score and the most consistent with the context lineage can be selected as the anchor point candidate for the matching item, so as to avoid the strong anchor point recognition result being affected by the unstable initial sorting of the candidate list.

[0063] Furthermore, the address resolution service selects the primary anchor point from deep to shallow according to the maximum level of the strong anchor point. If the later strong anchor point is an ancestor of the previous strong anchor point, or the previous strong anchor point is an ancestor of the later strong anchor point, then the two are merged into the same address path; if there is no common, ancestor or descendant relationship between the two strong anchor points, then candidate address paths are generated separately to handle the situation where there are multiple address fragments in the same text.

[0064] Furthermore, the address resolution service uses the main anchor point as a reference, collecting all ancestor matches upwards and all descendant matches downwards, and sorts them by administrative level to form candidate address paths. Paths with the same level and identical content are merged to avoid duplicate paths from entering subsequent scoring. Isolated matches not merged into the main path are discarded if they are not related, or if there are multiple candidates with a level greater than 3; the rest can be used to generate isolated single-node paths for subsequent NER fine-grained entity supplementation and low-confidence candidate output.

[0065] The following table shows the construction of a table for candidate address paths: The aforementioned candidate address path construction table illustrates that path construction is not simply a matter of concatenating administrative region keywords in text order. Instead, it uses strong anchor points as the center, collects nodes upwards and downwards along the standard administrative division lineage path, and forms a candidate address path set through path merging and isolated node processing.

[0066] In summary, S3 compresses multiple possible candidates generated in the coarse-ranking stage into a small number of candidate address paths with consistent lineage through feasibility pruning and path construction. Its core lies in using strong anchors and ancestor_bitmap for multi-candidate disambiguation, so that the candidate selection in scenarios with the same administrative region name, abbreviated administrative region name, and missing context no longer depends on simple string matching, but is constrained by the parent-child, ancestor, and descendant relationships of standard administrative divisions.

[0067] S4. Perform multi-dimensional scoring, path selection, and missing administrative level completion on candidate address paths. Specifically, step S4 of the method is as follows: comprehensively score the candidate address paths based on hierarchical completeness, hierarchical continuity, strong anchor point ratio, physical continuity, text coverage, and reverse order penalty, select address paths that meet the threshold, and complete the missing administrative levels.

[0068] Specifically, the address resolution service scores each candidate address path output by S3. The purpose of the scoring is to rank multiple feasible paths and provide interpretable credibility criteria for downstream services. Unlike schemes that only output the optimal solution, the path score in this embodiment is determined by multiple dimensions, including hierarchical completeness, hierarchical continuity, strong anchor ratio, physical continuity, text coverage, and inversion penalty.

[0069] In this embodiment, the path score can be expressed as: PathScore = 28 × Completeness + 22 × LevelContinuity + 22 × AnchorRatio + 16 × PhysicalContinuity + 12 × Coverage − 5 × InversionCount, where Completeness represents hierarchical integrity, LevelContinuity represents hierarchical continuity, AnchorRatio represents the proportion of strong anchors, PhysicalContinuity represents physical continuity, Coverage represents text coverage, and InversionCount represents the number of inversions of administrative levels after sorting by their original position. This formula reflects the combined relationship of five positive scoring dimensions and an inversion penalty item.

[0070] The following table shows the path scoring dimensions: The aforementioned path scoring dimension table illustrates that the path scoring in this embodiment is not a single string hit count score, but rather an evaluation that comprehensively considers path structure, administrative level, strong anchor points, text physical distance, and coverage. Specifically, `path_span_len` represents the span length of the actual text nodes in the candidate address path from the smallest `start` to the largest `end`; `avg_gap` represents the average character spacing between adjacent actual text nodes. In this embodiment, physical continuity and text coverage only count nodes that actually appear in the original text, excluding automatically completed nodes, to avoid artificially inflating path coverage or altering node spacing through completed nodes.

[0071] Furthermore, the number of levels to be covered in the hierarchy completeness is determined based on the administrative division structure. The typical administrative division path is based on four levels: provincial, municipal, district / county, and township / street. For the municipality model, the cross-level relationship from the provincial to the district / county level can be considered continuous; for the prefecture-level city-administered town model, the cross-level relationship from the municipal to the township / street level can be considered continuous. This allows for compatibility with municipalities such as Beijing, Shanghai, Tianjin, and Chongqing, as well as prefecture-level cities such as Dongguan, Zhongshan, Danzhou, and Jiayuguan-administered towns.

[0072] The table below shows the continuous rules for special administrative levels: The aforementioned special administrative level continuity rule table is used to explain that this embodiment does not require all address texts to explicitly show the four-level administrative level. Instead, it makes continuity judgments based on the standard administrative division tree and the equivalent mapping of specific levels, thereby reducing the risk of misjudgment of non-standard hierarchical structures such as municipalities and directly administered towns.

[0073] Furthermore, the address resolution service applies a reverse order penalty to candidate address paths. This penalty is calculated only for administrative region nodes that actually appear in the original text. After sorting the nodes according to their start positions in the original text, it checks whether there is a reverse arrangement of administrative levels from lower to higher levels. For example, if a county-level node appears before a city-level node in a path, and the text order does not conform to an acceptable inverted address representation, it can be counted as an inversion pair. Each inversion pair deducts 5 points. Automatically completed nodes, because their position=(-1,-1), do not have a true text order and are not subject to the reverse order penalty.

[0074] Furthermore, the address resolution service selects candidate address paths that meet a threshold after scoring. The threshold can be set according to the business scenario; for example, the threshold for automatic data entry with high confidence can be higher than the threshold for manual review.

[0075] After determining the candidate address paths, the address resolution service performs path completion. Path completion nodes are used to represent the structural integrity of the standard administrative division path, but are not considered as actual fragments appearing in the input text. PathCompleter locates the deepest administrative division node from the candidate address paths and finds the missing parent administrative division node based on the ancestor_bitmap of this deepest node, generating a completion node with position=(-1,-1).

[0076] Coverage calculation can be expressed as: Coverage = length(mergeIntervals({[start,end]|start>=0ANDend>=0})) / length(input_text); where mergeIntervals represents merging the intervals of the actual text nodes, start>=0ANDend>=0 is used to exclude autocomplete nodes with position=(-1,-1), and input_text represents the original input text.

[0077] The table below shows the participation rules for completing node calculations: The above-mentioned completion node calculation participation rule table is used to explain the boundaries of path completion. The completion node is not a matching entity in the original text, but a structured node derived from the ancestor relationship of standard administrative divisions; it should be marked as the source of autocomplete when output, and its position=(-1,-1) should not be used to participate in ordinary entity overlap deletion, nor should it be used to improve text coverage or physical continuity scores.

[0078] In summary, S4 transforms candidate address paths into sortable and interpretable scoring results through five-dimensional scoring and reverse order penalty; and through path completion based on ancestor_bitmap, addresses with default provincial, municipal, and district / county levels can still be normalized into complete standard administrative division paths. Therefore, this embodiment can maintain stable structured output for addresses such as Zhengzhou Jinshui in Henan, Chaoyang District in Beijing, and Humen Town in Dongguan, which omit suffixes, cross-level municipalities, and directly administered towns of prefecture-level cities.

[0079] S5. Identify fine-grained address entities in non-administrative regions using a sequence labeling model. Specifically, step S5 involves identifying fine-grained address entities in non-administrative regions using a sequence labeling model.

[0080] Specifically, administrative region entities are determined by the dictionary indexing, lineage trimming, path construction, path scoring, and path completion processes from S2 to S4, and are bound to the standard area_id. Fine-grained address entities such as roads, house numbers, neighborhoods or communities, buildings, units, floors, room numbers, POIs, sub-POIs, details, and villages are identified by the sequence labeling model. Therefore, administrative region normalization and fine-grained semantic generalization are handled by different mechanisms, avoiding the problem that pure NER models cannot output standard administrative region codes, and also avoiding the problem that pure dictionary rules are unable to identify detailed address fragments.

[0081] Specifically, the sequence labeling model can consist of a Chinese pre-trained encoder, a bidirectional long short-term memory network, a linear mapping layer, and a conditional random field layer. The encoder can use a Chinese pre-trained model to obtain character-level contextual semantic representations; BiLSTM is used to reinforce bidirectional sequence dependencies; the linear mapping layer outputs the emission scores of each label; the CRF layer uses the transition matrix to constrain the legality of the BIO label sequence, uses negative log-likelihood during the training phase, and uses Viterbi decoding during the inference phase.

[0082] Furthermore, the sequence labeling model can adopt the following structure: the encoder uses HFL / RBT3 or other pre-trained Chinese encoders to obtain character-level contextual semantic representations; the BiLSTM uses a two-layer bidirectional LSTM, and outputs contextual features after bidirectional concatenation; the linear mapping layer maps the contextual features to the label space; the CRF layer outputs the globally optimal BIO label sequence under label transfer constraints. In this embodiment, the specific pre-trained model, hidden layer dimensions, Dropout parameters, and optimizer configuration can be adjusted according to deployment resources and are not considered essential core limitations; the core lies in the hybrid division and merging of administrative region normalization paths and non-administrative region fine-grained NER results.

[0083] Furthermore, the tagging system can cover 15 entity types, generating 15 × 2 + 1 = 31 tags after BIO encoding. Here, O represents a non-entity character, BX represents the starting character of entity type X, and IX represents the internal characters of entity type X.

[0084] The following table shows the NER labeling system: The aforementioned NER label system table is for illustrative purposes. Although the model training labels can include administrative region labels such as PROVINCE, CITY, DISTRICT, and TOWN, the NER entities that ultimately participate in entity merging are mainly non-administrative region fine-grained address entities. In this embodiment, the sequence labeling model output is limited to non-administrative region fine-grained entities for merging, and administrative region entities are based on the normalized AddressPath result after fine-ranking.

[0085] Furthermore, during model inference, the tokenizer returns an offset map, which the address resolution service uses to map BIO tags back to the original character range. If an isolated IX tag appears, post-inference processing can fault-tolerantly treat it as the start of a new entity to avoid the loss of the entire entity due to local tag anomalies. For example, if a segment is labeled I-ROAD due to a model transfer error but the preceding character is not B-ROAD or I-ROAD, the I-ROAD can be corrected to B-ROAD before extraction continues.

[0086] The following table shows an example of BIO annotation: The above BIO annotation example table is used to illustrate that the model can recognize administrative district category labels and fine-grained labels. However, during the final merging, the administrative district category results do not directly use NER text fragments, but are generated by the AddressPath output by S4, so that the Beijing Municipality and Chaoyang District administrative district entities carry standard regional codes.

[0087] Furthermore, NER post-processing can include merging duplicate entities, substring determination, and suffix normalization. For example, entities with different suffixes, such as "Henan Province" and "Henan," can be merged using suffix normalization; for duplicate entities with complete inclusion relationships, more reasonable entities can be retained based on entity type, interval location, and confidence level. Post-processing does not change the standard area_id of the administrative region path; it is only used to clean up duplicate, fragmented, or type-conflicting results in fine-grained NER output.

[0088] In summary, S5 supplements non-administrative region fine-grained address entities such as roads, house numbers, residential areas, buildings, units, room numbers, and POIs with a sequence labeling model. This enables this embodiment to possess both the ability to normalize according to administrative division standards and the ability to semantically recognize detailed addresses in free text. Administrative region entities do not rely on NER secondary matching but are determined by indexing and path fine-grained ranking; NER focuses on fine-grained address fragments outside the standard administrative division system, and the two have a clear division of labor.

[0089] S6. Merge the normalized administrative region entities and fine-grained entities, and output the structured address resolution result. Specifically, step S6 is to merge the normalized administrative region entities and fine-grained address entities generated by the address path according to their original positions, and output the structured address resolution result.

[0090] Specifically, after completing the candidate address path scoring and completion in S4 and the fine-grained NER identification in S5, the address resolution service enters the resolution scheduling and entity merging stage. In this stage, link IDs can be generated first, followed by obtaining coarse-ranked matching results, fine-ranked path results, and fine-grained NER entity results in sequence, and merging administrative region entities and non-administrative region entities into a unified entity list.

[0091] Furthermore, administrative region entities are primarily generated from the refined AddressPath, with hierarchical labels including Province, City, District, and Town, each corresponding to a standard administrative division node. Each normalized administrative region entity includes at least a standard name, area code (area_id), administrative level, original text position (position), and an indicator for whether auto-completion is enabled. For completed nodes generated by S4, their position is marked as (-1, -1), indicating that the node does not actually appear in the original text but is a standard path node derived based on ancestor_bitmap.

[0092] Furthermore, the administrative region type in the original NER output is skipped during the entity conversion stage to avoid duplication or conflict with the normalized AddressPath administrative region entity. NER mainly handles the road, house number, community, neighborhood, building, unit, floor, room number, POI, sub-POI, and details fields. During merging, the address resolution service first performs NER entity deduplication, cleaning, and type-level filtering, then concatenates the administrative region entity with the NER entity, and sorts them in ascending order according to the original start position; when the region label overlaps with the actual location of non-region labels such as roads, POIs, and details, the region label is retained first.

[0093] The following table shows the entity merging rules: The entity merging rule table described above illustrates that the entity merging in this embodiment is not a simple concatenation of NER results, but rather primarily uses the normalized administrative region entities generated by fine-grained AddressPath sorting, supplemented by the fine-grained address entities identified by NER. When the POI, road, or detail entities identified by NER overlap with the administrative region entities, if the administrative region entity is a standard administrative region node that actually matches the original text, the administrative region entity is retained first; if the administrative region node is a complete node with position=(-1,-1), it is not involved in ordinary overlap deletion; if the administrative region hit originates from a business alias mapping and the fragment also has POI semantics, the POI entity should be retained and the mapping source should be marked in the administrative region entity to avoid losing the true fine-grained address information.

[0094] Furthermore, the address resolution service can output the final result as a structured address resolution result. The structured address resolution result can include fields such as the original input text, standard administrative region path, normalized administrative region entity list, non-administrative region fine-grained entity list, merged unified entity list, path score, whether completion is required, low confidence indicator, and link ID.

[0095] The following table shows the fields of the structured address resolution result: The structured address parsing result field table mentioned above is used to explain that the final output result not only includes the segmented address entity, but also the standard administrative division code, path score and completion status. Therefore, downstream logistics distribution, customer address cleaning, LBS positioning, government population management and other businesses can directly use the standard area_id for normalization and database entry, or manually review or process low confidence addresses according to path_score and confidence_flag.

[0096] For example, given the input text "No. 19, Sanlitun Road, Chaoyang District, Beijing", S2 recalls matching items such as administrative districts such as Chaoyang District, Beijing; S3 constructs candidate paths from Beijing to Chaoyang District; S4 calculates the path score and completes the parent node when necessary; S5 identifies Sanlitun Road as ROAD and No. 19 as ROADNUM; and S6 finally outputs City and District entities with standard area codes, as well as fine-grained entities such as ROAD and ROADNUM.

[0097] In summary, S6 achieves a fusion output of administrative region standardization and detailed address semantic recognition by uniformly sorting, conflict handling, and threshold filtering the normalized administrative region entities generated by AddressPath and the fine-grained address entities identified by NER. This merging method ensures that administrative region normalization is guaranteed by dictionary indexing and path fine-grained sorting, while fine-grained address supplementation is guaranteed by NER. The two have a clear division of labor and are output through a unified entity list.

[0098] In any of the above embodiments, the method may further include non-stop hot updates. Specifically, it involves constructing a new in-memory data connection, a two-level in-memory index, an AC automaton, a coarse-ranking matcher, and a fine-ranking decision maker based on the new version of the standard administrative division data; if the construction is successful, the runtime references are replaced once; if the construction fails, the old version is maintained and services continue.

[0099] Specifically, the address resolution service can receive hot update requests via the reload interface or other triggering methods. Upon receiving a request, the system reads the current versioned data configuration and retrieves the new administrative division data from the remote data table. After reading, the new data is registered in a new in-memory data connection, and indexes such as parent_id and floor are created. Subsequently, a new IndexManager is built based on the new in-memory data connection, all administrative division nodes are loaded, and GlobalHashIndex, HierarchicalTrie, path_code, and ancestor_bitmap are generated. Then, a new CoarseMatcher is built based on the new indexes, and the AC automaton is regenerated. Finally, a new FineRanker is built based on the new Trie, which includes a new feasibility trimmer, path builder, path scorer, and path completer.

[0100] The following table shows the hot update process without system downtime: The above-described non-downtime hot update process table illustrates that, in this embodiment, the hot update does not modify nodes, keywords, or bitmaps item by item within the old index. Instead, it adopts a method of first building the new state and then replacing the runtime references. In this embodiment, the replacement action should be based on the overall reference replacement of the runtime state object, and fields should not be modified locally within the old IndexManager, CoarseMatcher, or FineRanker to avoid concurrent requests reading the half-built state.

[0101] Furthermore, once the new state is successfully constructed, the address resolution service replaces the runtime references all at once, ensuring that resolution requests entering after the replacement use the new state. Requests that have already entered the resolution process before the replacement still hold their old object references and are unaffected by the new state construction process. After the replacement is complete, the system can attempt to close the old memory data connection and release the old data state; if any step of new data loading, index construction, matcher construction, or fine sorter construction fails, reference replacement will not be performed, and the old state will continue to provide services.

[0102] Furthermore, if the address resolution service is deployed using multiple processes or multiple replicas, runtime reference replacement within a single process is only effective for that process. In multi-process or multi-replica scenarios, each worker process should perform hot updates separately, or an external orchestration system should trigger rolling loading to ensure that each service instance ultimately loads a consistent data version. This processing does not change the hot update mechanism of building the new state first and then replacing the references as a whole in this embodiment; it only performs engineering corrections to ensure data consistency in multi-instance deployment scenarios.

[0103] In summary, the supplementary implementation method achieves non-stop hot updates of administrative division data by replacing the entire runtime state. This mechanism avoids direct modification of the internal structure of the old index during the hot update process, reducing the risk of concurrent requests reading from a half-built state; at the same time, it maintains the old state to continue service when the build fails, ensuring the availability of the address resolution service during the administrative division data update process.

[0104] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.< / int> < / areanode>

Claims

1. A hybrid Chinese address resolution method based on sequence labeling and double-layer indexing, characterized in that: Performed by the processor of the address resolution service, including: S1. Load standard administrative division data and construct a two-layer memory index, which includes a global inverted hash index that maps candidate administrative division nodes with keywords, and a hierarchical Trie tree that organizes nodes according to administrative parent-child levels and pre-computes the ancestor bitmap. S2. Receive the Chinese address text to be parsed, and recall keyword matching items and their candidate administrative region nodes on the cleaned view that retains the original offset mapping through the AC automaton; S3. Use the ancestor bitmap to determine the same, ancestor or descendant relationship between candidate nodes, and suppress the included matching items only when the text range contains and the relationship exists; S4. Based on candidate uniqueness, candidate node level and neighboring ancestor matching, determine strong anchor points, and prune multiple candidate nodes according to strong anchor points and ancestor bitmaps, and construct candidate address paths along the same lineage path; S5. The candidate address paths are comprehensively scored based on hierarchical completeness, hierarchical continuity, strong anchor point ratio, physical continuity, text coverage, and reverse order penalty. The address paths that meet the threshold are selected and the missing administrative levels are filled in. S6. Use a sequence labeling model to identify fine-grained address entities outside administrative regions, merge the normalized administrative region entities generated by the address path with the fine-grained address entities according to their original positions, and output the structured address parsing results.

2. The method of claim 1, wherein: The administrative division node includes a region code, standard name, abbreviation, alias, level, and parent node code; the standard name, abbreviation, and alias are all written as keywords into the global inverted hash index, and the ancestor bitmap is generated by superimposing the parent node's ancestor bitmap with the parent node's identifier.

3. The method of claim 1, wherein: The cleaned view filters out numbers, punctuation, and special characters while retaining Chinese and English letters, while the original text view retains the complete text; the matching position of the AC automaton on the cleaned view is backfilled to the original text position through offset mapping.

4. The method of claim 1, wherein: The suppression of included matches includes: sorting matches in ascending order of start position and descending order of end position; deleting the short match when a short match is included by a long match, and the candidate nodes of the two are the same or in the same lineage path, otherwise retaining it.

5. The method of claim 1, wherein: The pruning of multiple candidate nodes includes: collecting candidate nodes and their ancestors to form a set of context nodes, and using the unique candidate as a strong anchor point; for multiple candidates, priority is given to retaining candidates that have a strong anchor point as an ancestor or are an ancestor of a strong anchor point, and if necessary, continuing to converge according to the deepest strong anchor point.

6. The method of claim 5, wherein: The pruning of multiple candidate nodes also includes fragment deduplication, position pruning, and prefix chain pruning; fragment deduplication uses rolling hashing, position pruning selects the best candidates based on lineage support, keyword length, candidate uniqueness, and hierarchical information content, and prefix chain pruning removes candidates that have no lineage relationship with the confirmed upper-level nodes.

7. The method of claim 1, wherein: The construction of candidate address paths includes: determining strong anchor point scores based on candidate level not lower than the city level, candidate uniqueness, and nearby ancestor matching; using the strongest anchor point at the highest level as the main anchor point, collecting its ancestor matching items and descendant matching items, and sorting them by administrative level to form candidate address paths.

8. The method according to claim 1, characterized in that: The hierarchical continuity is considered as continuous from the provincial level to the district / county level under the municipality model, and as continuous from the city level to the township / street level under the prefecture-level city-administered town model. The completion node is determined by the ancestor bitmap of the deepest node and marked as (-1, -1). It does not participate in the calculation of text coverage, physical continuity and reverse order penalty.

9. The method according to claim 1, characterized in that: The sequence labeling model includes a Chinese pre-trained encoder, a bidirectional long short-term memory network, a linear mapping layer, and a conditional random field layer; it identifies entities such as roads, house numbers, residential areas or communities, buildings, units, floors, room numbers, points of interest, sub-points of interest, details, and village groups. When merging, it skips the administrative region entities it outputs and prioritizes retaining normalized administrative region entities.

10. The method according to claim 1, characterized in that: It also includes non-stop hot updates: building new in-memory data connections, two-level in-memory indexes, AC automata, coarse-ranking matchers, and fine-ranking decision makers based on the new standard administrative division data; replacing runtime references once the build is successful, and maintaining the old version to continue service if the build fails.