A Chinese address administrative division standardization method, system and device
By constructing an administrative division mapping relationship model and word segmentation algorithm, the irregularity of Chinese address data is solved, the standardized completion of address data is achieved, and its processing efficiency and application value in computers is improved.
Patent Information
- Application Number
- CN202310450187.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-24
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2043-04-24
AI Technical Summary
Due to the ambiguity of connotation and diversity of forms, Chinese address data has complex forms and is difficult for computers to understand and process, and has problems such as incomplete elements, inconsistent expressions, and easy to cause ambiguity, which affects its liquidity and potential value.
By constructing an address element pair and mapping relationship model corresponding to administrative divisions, combining the state transfer mode and word segmentation algorithm of the logical dynamic system, the standardized completion of Chinese addresses is realized and transformed into a standard address structure containing complete administrative divisions.
It improves the overall quality and utilization value of Chinese address data, realizes complete storage, precise matching and standard processing in computers, and supports the integration and analysis of address data.
Smart Images

Figure CN116401331B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data processing, and in particular, to a method, system and device for standardizing administrative divisions of Chinese addresses. Background Art
[0002] A Chinese address is a short-text natural language string composed of multiple address element entities combined according to a certain sorting rule to describe spatial location information. As an important type of information that can associate different data sources, Chinese addresses have become important basic spatial data in various economic activities and important carriers for information transmission of various government and enterprise affairs. At the same time, they have also penetrated into many aspects of personal life.
[0003] Due to the polysemy of Chinese connotations and the diversity of forms, Chinese addresses composed of Chinese and special characters not only contain place name information but may also contain non-standard descriptions of spatial information. This makes Chinese addresses, as a non-standard natural language string and an unstructured descriptive data, have problems such as complex and diverse forms and difficulty for computers to understand and process. In addition, due to different collection methods, inconsistent recording methods, and inconsistent naming standards of address elements for address data, there are problems such as incomplete elements, inconsistent expressions, and easy to cause ambiguity in Chinese address data. These problems greatly affect the overall quality of Chinese address data, making it unable to be directly used for matching, statistics, and analysis. Eventually, it will affect the circulation and potential value of address data, hinder the research and application of address data, and be difficult to meet the needs of the development of informatization. Summary of the Invention
[0004] To solve the above problems, the present application proposes a method for standardizing administrative divisions of Chinese addresses, including:
[0005] Construct address element pairs corresponding to the administrative division according to the administrative division and the index number corresponding to the administrative division, and obtain address element sets corresponding to administrative divisions at all levels according to the address element pairs;
[0006] Construct a state space corresponding to the address element set, determine the hierarchical subordination relationship between adjacent-level administrative divisions, and determine a mapping relationship model between adjacent-level administrative divisions according to the hierarchical subordination relationship;
[0007] Determine a hierarchical association matrix corresponding to the administrative division according to the mapping relationship model;
[0008] Obtain the original address characters of the administrative division to be standardized, and segment the original address characters to obtain the original address structure corresponding to the original address characters;
[0009] Obtain the missing completion conditions corresponding to the administrative division, and determine whether the original address structure satisfies the missing completion conditions. If so, update the original address structure according to the hierarchical association matrix to obtain the standard address structure corresponding to the administrative division to be standardized.
[0010] In an implementation manner of the present application, according to the address element pairs, obtain the address element sets corresponding to administrative divisions at all levels, specifically including:
[0011] According to the address element pairs, obtain the single-level address element sets corresponding to administrative divisions at all levels respectively;
[0012] For each single-level address element set, determine the number of address elements in the single-level address element set, and according to the number of address elements, determine the index number range of the administrative division corresponding to the single-level address element set;
[0013] According to the index number range, index the administrative divisions corresponding to the single-level address element sets to obtain the address element sets corresponding to administrative divisions at all levels.
[0014] In an implementation manner of the present application, construct the state space corresponding to the address element set, specifically including:
[0015] Determine the index numbers corresponding to each administrative division in the address element set, and according to the index numbers, establish logical vectors equivalent to the index numbers;
[0016] According to the logical vectors, construct the state space corresponding to the address element set.
[0017] In an implementation manner of the present application, according to the hierarchical subordination relationship, determine the mapping relationship model between adjacent-level administrative divisions, specifically including:
[0018] For the state space corresponding to the address element sets of administrative divisions at all levels, divide the state space into several state sub-spaces;
[0019] According to the hierarchical subordination relationship, determine the first mapping relationship between the state sub-space and the state space corresponding to its adjacent-level administrative division, so as to determine the second mapping relationship between adjacent-level administrative divisions according to the first mapping relationship;
[0020] According to the second mapping relationship, construct the mapping relationship model between adjacent-level administrative divisions.
[0021] In an implementation manner of the present application, according to the mapping relationship model, determine the hierarchical association matrix corresponding to the administrative division, specifically including:
[0022] Determine the weak hierarchical association matrix corresponding to the administrative division according to the mapping relationship model;
[0023] Determine whether a specified administrative division in the administrative division belongs to the adjacent-level administrative division corresponding to the administrative division according to the weak hierarchical association matrix;
[0024] If so, perform iterative calculation on the weak hierarchical association matrix to obtain the hierarchical association matrix corresponding to the administrative division.
[0025] In an implementation manner of the present application, segment the original address characters to obtain the original address structure corresponding to the original address characters, which specifically includes:
[0026] Segment the original address characters to obtain an address element sequence composed of multiple address elements corresponding to the original address characters, and sequentially match the multiple address elements with the address element set;
[0027] In the case where there is a matching address element set, obtain the specified address element in the address element set that matches the address element, and add the specified address element to the original address structure of the original address characters;
[0028] In the case where there is no matching address element set, obtain the original address structure corresponding to the original address characters according to other address elements after the current address element in the address element sequence.
[0029] In an implementation manner of the present application, before updating the original address structure according to the hierarchical association matrix, the method further includes:
[0030] Obtain the target value that meets the missing complement condition and the target address element corresponding to the target value;
[0031] Determine the address element set where the target address element is located, and generate the logical matrix corresponding to the administrative division according to the index number in the address element set.
[0032] In an implementation manner of the present application, update the original address structure according to the hierarchical association matrix to obtain the standard address structure corresponding to the administrative division to be standardized, which specifically includes:
[0033] Construct the number complement matrix corresponding to the administrative division according to the hierarchical association matrix and the logical matrix; wherein, the number complement matrix is composed of the index numbers of the administrative divisions to which the address elements belong in their corresponding address element sets;
[0034] Obtain the address element pairs corresponding to each element in the numbered completion matrix to construct the administrative division completion matrix corresponding to the administrative division;
[0035] Update the original address structure according to the administrative division completion matrix to obtain the standard address structure corresponding to the administrative division to be standardized.
[0036] An embodiment of the present application provides a Chinese address administrative division standardization system, and the system includes:
[0037] An administrative division element matching module, configured to construct an address element pair corresponding to the administrative division according to the administrative division and the index number corresponding to the administrative division, and obtain an address element set corresponding to administrative divisions at all levels according to the address element pair;
[0038] An administrative division level association module, configured to construct a state space corresponding to the address element set, determine the hierarchical subordination relationship between adjacent-level administrative divisions, and determine a mapping relationship model between adjacent-level administrative divisions according to the hierarchical subordination relationship;
[0039] It is also configured to determine a hierarchical association matrix corresponding to the administrative division according to the mapping relationship model;
[0040] An administrative division identification and conversion module, configured to obtain the original address characters of the administrative division to be standardized, perform word segmentation on the original address characters to obtain an original address structure corresponding to the original address characters;
[0041] An administrative division standard completion module, configured to obtain a missing completion condition corresponding to the administrative division, determine whether the original address structure satisfies the missing completion condition, and if so, update the original address structure according to the hierarchical association matrix to obtain a standard address structure corresponding to the administrative division to be standardized.
[0042] An embodiment of the present application provides a Chinese address administrative division standardization device, and the device includes:
[0043] At least one processor;
[0044] And a memory communicatively connected to the at least one processor;
[0045] Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can:
[0046] Construct an address element pair corresponding to the administrative division according to the administrative division and the index number corresponding to the administrative division, and obtain an address element set corresponding to administrative divisions at all levels according to the address element pair;
[0047] Construct the state space corresponding to the set of address elements, determine the hierarchical subordination relationship corresponding to adjacent-level administrative divisions, and according to the hierarchical subordination relationship, determine the mapping relationship model between adjacent-level administrative divisions;
[0048] According to the mapping relationship model, determine the hierarchical association matrix corresponding to the administrative division;
[0049] Obtain the original address characters of the administrative division to be standardized, and segment the original address characters to obtain the original address structure corresponding to the original address characters;
[0050] Obtain the missing complement condition corresponding to the administrative division, determine whether the original address structure satisfies the missing complement condition, and if so, update the original address structure according to the hierarchical association matrix to obtain the standard address structure corresponding to the administrative division to be standardized.
[0051] A Chinese address administrative division standardization method proposed by this application can bring the following beneficial effects:
[0052] Combined with the state transition mode of the logical dynamic system and the word segmentation algorithm, a matching association model and related implementation algorithms for Chinese address administrative division standardization and complementation are proposed. By converting the unstructured Chinese address described by using non-standard natural language strings to describe location information into a standard address structure including complete administrative divisions, it is possible to solve the problems of incorrect arrangement of administrative division subordination relationships, omission of administrative division feature words, and missing of some administrative divisions existing in the Chinese address expression. Brief Description of the Drawings
[0053] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0054] Figure 1 It is a schematic flowchart of a Chinese address administrative division standardization method provided by an embodiment of the present application;
[0055] Figure 2 It is a schematic architecture diagram of a Chinese address administrative division standardization system provided by an embodiment of the present application;
[0056] Figure 3 It is a schematic structural diagram of a Chinese address administrative division standardization device provided by an embodiment of the present application. Detailed Description of the Embodiments
[0057] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.
[0058] As the key to address data quantification, Chinese address standardization is to convert a large number of unstructured addresses described by natural language strings into accurate and standardized Chinese address structures. Its purpose is to facilitate the representation, storage, matching, processing and application of address data in computers, so as to fully and effectively aggregate relevant data with standardized address structures as the core, and then realize the integration and analysis of relevant information and provide decision-making and services supported by standard address data. In the study of address standardization, the standardization and completion of administrative divisions is a technical problem that must be solved. A standard Chinese address should record and express administrative divisions at all levels (provincial administrative divisions / prefectural administrative divisions / county administrative divisions / township administrative divisions / village administrative divisions) in a complete and orderly manner, so as to effectively match the geographical location in the real world. However, in real life, when people record and express addresses, they usually encounter problems such as incorrect arrangement of administrative division affiliation, omission of administrative division feature words, and omission of some administrative divisions. These problems will make the Chinese address incomplete and information missing, which will affect the quality and value of the Chinese address. The standardized completion of Chinese address administrative divisions is to solve the problem of irregular administrative divisions, complete the incomplete address data into complete and valid address data, and improve the overall quality and utilization value of the address data.
[0059] The technical solutions provided by various embodiments of the present application are described in detail below in conjunction with the accompanying drawings.
[0060] like Figure 1 As shown, a method for standardizing Chinese address administrative divisions provided in an embodiment of the present application includes:
[0061] S101: constructing address element pairs corresponding to administrative divisions according to administrative divisions and index numbers corresponding to the administrative divisions, and obtaining address element sets corresponding to administrative divisions at various levels according to the address element pairs.
[0062] Administrative divisions are regions divided into different levels for the convenience of administrative management, as shown in Table 1:
[0063]
[0064] As shown in Table 1, administrative divisions are divided into five different levels, namely provincial level, prefecture level, county level, township level, and village level. Among them, the names of the first-level administrative divisions and the second-level administrative divisions are unique among administrative divisions of the same level, and the specific names of all the next-level administrative divisions belonging to the same administrative division are not repeated.
[0065] It can be seen from this that a valid original address record should at least contain one of the above five levels of administrative divisions, as well as other detailed location information other than administrative divisions. Otherwise, the address may not be able to uniquely and accurately correspond to the spatial location in the real world.
[0066] In order to facilitate the index matching of each level of administrative division, it is necessary to construct an associated form of administrative division index. Given an administrative division a, design a corresponding mechanism to associate it with a unique numerical index number n, and construct a binary address element pair (a, n) to achieve the equivalent conversion between the administrative division and its index number.
[0067] After constructing the address element pairs, it is necessary to obtain the single-level address element sets corresponding to each level of administrative division. For example, for the first-level address element set, the corresponding single-level address element set can be expressed as:
[0068]
[0069] where α1 is the number of first-level administrative divisions, represents the first-level administrative division corresponding to the index number i, and the administrative divisions corresponding to different index numbers are different, that is, when 1 ≤ i1 ≠ i2 ≤ α1 holds and represent different administrative divisions respectively. Further, let the parameter k of the administrative division level be initially set to 2.
[0070] For each single-level address element set, according to the subordination relationship of adjacent-level administrative divisions, for each address element pair it is possible to query and obtain all the sets composed of the k-level administrative divisions belonging to Determine the number of address elements in the single-level address element set According to the number of address elements, determine the index number interval of the administrative division corresponding to the single-level address element set. The index number interval is Through the integer τ(d) in this interval, the administrative division corresponding to the single-level address element set can be indexed to obtain the address element sets corresponding to each level of administrative division. Let i = 1, 2,..., α in turn k-1 , the corresponding index numbers can be matched for all k-level administrative divisions, and thus the k-level address element set can be obtained:
[0071]
[0072] Among them, represents the administrative division at the k-th level corresponding to the index number i, and satisfies
[0073] Finally, the set of address elements corresponding to administrative divisions at all levels can be obtained
[0074] S102: Construct the state space corresponding to the set of address elements, determine the hierarchical membership relationship between adjacent-level administrative divisions, and determine the mapping relationship model between adjacent-level administrative divisions according to the hierarchical membership relationship.
[0075] For in the set of address elements, determine the index numbers 1 ≤ i ≤ α1 corresponding to each administrative division . According to this index number, a logical vector equivalent to it can be established where represents the i-th column of the α1-order identity matrix . Based on this, according to the logical vector, the state space corresponding to the set of address elements can be constructed Furthermore, let the parameter k of the administrative division level be initially set to 2.
[0076] For the set of address elements The same method as above can be used to construct the equivalent k-level state space where represents the i-th column of the α k -order identity matrix .
[0077] According to the meaning of the set of administrative division address elements mentioned above, for the state space corresponding to administrative divisions at all levels, the state space can be divided into several k-level state subspaces. The state subspace can be expressed as According to the hierarchical membership relationship between adjacent-level administrative divisions, the set of state subspaces and the state space corresponding to its adjacent-level administrative division
[0078]
[0079] Furthermore, according to the first mapping relationship, the second mapping relationship between adjacent-level administrative divisions and can be determined. The second mapping relationship is as follows:
[0080]
[0081] Among them,
[0082] From the perspective of the state transition of the logical dynamic system, the hierarchical subordination relationship between the administrative division at the k-th level numbered i and the administrative division at the (k - 1)-th level numbered j is equivalent to the logical vector transferring to after one step. Thus, according to the second mapping relationship, a and can be established. The mapping relationship model is as follows:
[0083]
[0084] Among them, and represent the state variables at the (k - 1)-th and k-th levels respectively, that is, they respectively represent the logical vectors corresponding to the index numbers of the address elements in and equivalent and corresponding, represents the Kronecker product, represents a -dimensional row vector with all elements being 1, where i = 1, 2,..., α k-1 .
[0085] S103: Determine the hierarchical association matrix corresponding to the administrative division according to the mapping relationship model.
[0086] After obtaining the mapping relationship model, the k-th level weak hierarchical association matrix corresponding to the administrative division can be determined accordingly. The weak hierarchical association matrix can be expressed as:
[0087]
[0088] The weak hierarchical association matrix integrates and quantifies the hierarchical subordination relationship between adjacent-level administrative divisions, and can realize the matching association between the current administrative division and the upper-level administrative division. That is to say, according to the weak hierarchical association matrix, it is determined whether the specified administrative division in the administrative division belongs to the adjacent-level administrative division corresponding to the administrative division, which is manifested as if the i-th column of is
[0089] then it indicates that the administrative division at the k-th level numbered i belongs to the administrative division at the (k - 1)-th level numbered j.
[0090]
[0091] Among them, F kThe initial value is F1 = [1, 2, …, α1]. According to the above calculation principle, it can be known that the hierarchical association matrix integration includes the hierarchical subordination relationships of multiple administrative regions at different levels. The index numbers of higher-level administrative regions that are complemented can be obtained based on the index numbers of administrative regions, realizing the matching and association between the current administrative region and the previous administrative regions at all levels. The complementing principle is: If F k The i-th column of k-1 is [i1, i2, …, i T , i] , it indicates that the k-level administrative region with the index number i
[0092] Repeating the above process, the final sequence of hierarchical association matrices {F2, F3, F4, F5} can be obtained. Each hierarchical association matrix in this sequence is respectively matched with its corresponding administrative region level.
[0093] S104: Obtain the original address characters of the administrative region to be standardized, and segment the original address characters to obtain the original address structure corresponding to the original address characters.
[0094] First, obtain the original address characters of the administrative region to be standardized, and use the BiLSTM+CRF word segmentation algorithm to segment each original address character A in to obtain an address element sequence composed of multiple address elements corresponding to the original address characters. The address element sequence can be expressed as
[0095] Combined with the preset standard address form, the address structure can be defined as:
[0096]
[0097] It is necessary to initialize the above address structure, that is, let ξ i = 0, i = 1, 2, …, 6.
[0098] Then, match the multiple address elements with the address element set in turn. In the case where there is a matching address element set, obtain the specified address element d matching the address element in the address element set k , administrative level and the set composed of its corresponding index number in , and add the specified address element d to the original address structure of the above original address characters, and update k in to d to d k, γ = 1. In the case where there is no set of matching address elements, the original address structure corresponding to the original address characters is obtained according to the other address elements in the address element sequence that are after the current address element After that The ξ6 in is updated to γ = 0.
[0099] If γ = 1, let k = k + 1 and continue to execute the above process. If γ = 0, output the original address structure and the set corresponding to the non-zero element d therein
[0100] S105: Obtain the missing complement condition corresponding to the administrative division, determine whether the original address structure meets the missing complement condition. If so, update the original address structure according to the hierarchical association matrix to obtain the standard address structure corresponding to the administrative division to be standardized.
[0101] Combined with the original address structure obtained in S104 The missing complement condition corresponding to the administrative division can be obtained, which is specifically as follows:
[0102]
[0103] Determine whether the original address structure meets the above missing and incomplete conditions. If not, the current administrative division cannot be complemented or does not need to be complemented, and becomes the marked address structure after standardization.
[0104] If it meets the conditions, it is necessary to update the original address structure according to the hierarchical association matrix to obtain the standard address structure corresponding to the administrative division to be standardized.
[0105] Specifically, obtain the target value that meets the missing complement condition and the target address element corresponding to the target value. Among them, the target value is the largest number of elements that meet the missing complement condition in, 2 ≤ θ ≤ 5, and the target address element is ξ θ . After determining the target address element, the address element set where it is located can be determined Furthermore, according to the index number n in the address element set i , generate the logical matrix corresponding to the administrative division
[0106] Furthermore, according to the hierarchical association matrix and the logical matrix, construct the number complement matrix corresponding to the administrative division. The number complement matrix can be expressed as:
[0107]
[0108] Among them, h i,j represents the index number of the administrative division to which the address element pair (ξ θ , n i ) belongs in its corresponding address element set , where k = 1, 2,..., θ - 1.
[0109] Furthermore, obtain the address element pairs corresponding to each element h i,k in the k-th column of the above-mentioned number-completed matrix H to construct the administrative division completion matrix corresponding to the administrative division. The administrative division completion matrix is expressed as:
[0110]
[0111] Among them, d i,k represents the k-level administrative division completed based on the administrative division ξ θ , where k = 1, 2,..., θ - 1. Each row in the matrix D corresponds to a permutation of the administrative divisions completed based on the administrative division ξ θ . According to this administrative division completion matrix, the original address structure can be updated to obtain the standard address structure corresponding to the administrative division to be standardized.
[0112] Specifically, if s = 1, then for 1 ≤ i < θ, the ξ in the original address structure i needs to be updated to d 1,i , so as to obtain the corresponding standard address structure. Otherwise, it is necessary to determine the row vector in the administrative division completion matrix D with the highest coincidence rate with the elements in the original address structure . For the row vector that meets the conditions, according to the situation when s = 1, the original address structure is updated to obtain the final standard address structure.
[0113] The present invention combines the state transition mode of the logical dynamic system and the BiLSTM+CRF algorithm to propose a matching and correlation model and related implementation algorithms for standardizing and completing the administrative divisions of Chinese addresses. By converting unstructured Chinese addresses described using non-standard natural language strings to standard address structures containing complete administrative divisions, it can solve problems such as incorrect permutations of administrative division subordination relationships, omission of administrative division characteristic words, and missing of some administrative divisions in Chinese address expressions, and can improve the overall quality of Chinese address data for its complete storage, accurate matching, standard processing, and wide application in computers, and can provide effective support for the integration and intercommunication of address data in different scenarios and a solid foundation for further exploring greater potential value of address data.
[0114] The above is the method embodiment proposed in this application. Based on the same idea, some embodiments of this application also provide a system and a device corresponding to the above method.
[0115] An embodiment of this application provides a Chinese address administrative division standardization system, which includes:
[0116] An administrative division element matching module, configured to construct an address element pair corresponding to an administrative division according to the administrative division and the index number corresponding to the administrative division, and obtain an address element set corresponding to administrative divisions at all levels according to the address element pair;
[0117] An administrative division hierarchy association module, configured to construct a state space corresponding to the address element set, determine the hierarchical subordination relationship between adjacent-level administrative divisions, and determine a mapping relationship model between adjacent-level administrative divisions according to the hierarchical subordination relationship;
[0118] It is also configured to determine a hierarchical association matrix corresponding to the administrative division according to the mapping relationship model;
[0119] An administrative division identification and conversion module, configured to obtain the original address characters of the administrative division to be standardized, perform word segmentation on the original address characters to obtain an original address structure corresponding to the original address characters;
[0120] An administrative division standard complement module, configured to obtain a missing complement condition corresponding to the administrative division, determine whether the original address structure meets the missing complement condition, and if so, update the original address structure according to the hierarchical association matrix to obtain a standard address structure corresponding to the administrative division to be standardized.
[0121] Figure 2 It is a schematic diagram of the architecture of a Chinese address administrative division standardization system provided by an embodiment of this application. As Figure 2 shown, the administrative division identification and conversion module can perform word segmentation on the original address characters through a word segmentation algorithm, and process the segmented original address characters to obtain an original address structure.
[0122] The administrative division element matching module can establish an index association form for each level of administrative division to achieve index matching of the same-level administrative divisions. This module mainly numbers each level of administrative division regularly based on the inherent characteristics and internal associations of the administrative division, and then generates an address element pair composed of the administrative division and the corresponding index number, and finally calculates the address element set of each level of administrative division.
[0123] The administrative division level association module is used to establish a hierarchical association model for multi-level administrative divisions and achieve the matching association of administrative divisions at different levels. This module mainly models and represents the hierarchical association of adjacent-level administrative divisions in combination with the state transition mode of the logical dynamic system, and then constructs a weak hierarchical association matrix for the corresponding level. Finally, through the iterative calculation of the induction method, a hierarchical association matrix that can deeply describe the external characteristics and internal laws of administrative divisions at different levels is obtained.
[0124] The administrative division standard complement module is used to establish a standardized complement algorithm for administrative divisions and achieve the construction of a standard address structure. This module mainly uses the condition for judging the complement of missing administrative divisions proposed based on the original address structure to determine the feasibility of completely complementing the missing administrative divisions. When the missing parts can be complemented, the corresponding hierarchical association matrix is used to calculate the administrative division complement matrix, and finally the standard address structure is determined in combination with the original address structure.
[0125] Figure 3 This is a schematic structural diagram of a Chinese address administrative division standardization device provided by an embodiment of the present application. As Figure 3 shown, it includes:
[0126] At least one processor; and,
[0127] A memory communicatively connected to at least one processor; wherein,
[0128] The memory stores instructions executable by at least one processor. The instructions are executed by at least one processor so that at least one processor can:
[0129] Construct an address element pair corresponding to the administrative division according to the administrative division and the index number corresponding to the administrative division, and obtain an address element set corresponding to each level of administrative division according to the address element pair;
[0130] Construct a state space corresponding to the address element set, determine the hierarchical membership relationship corresponding to adjacent-level administrative divisions, and determine a mapping relationship model between adjacent-level administrative divisions according to the hierarchical membership relationship;
[0131] Determine a hierarchical association matrix corresponding to the administrative division according to the mapping relationship model;
[0132] Obtain the original address characters of the administrative division to be standardized, segment the original address characters to obtain the original address structure corresponding to the original address characters;
[0133] Obtain the missing complement condition corresponding to the administrative division, determine whether the original address structure satisfies the missing complement condition. If so, update the original address structure according to the hierarchical association matrix to obtain the standard address structure corresponding to the administrative division to be standardized.
[0134] The embodiments in the present application are all described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the relevant content.
[0135] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, commodity or device including the said element.
[0136] The above description is only for the embodiments of the present application and is not intended to limit the present application. For those skilled in the art, various modifications and changes can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A method for standardizing administrative divisions of Chinese addresses, characterized in that, The method includes: Construct address element pairs corresponding to the administrative division according to the administrative division and the index number corresponding to the administrative division, and obtain address element sets corresponding to administrative divisions at all levels according to the address element pairs; Construct a state space corresponding to the address element set, determine the hierarchical subordination relationship between adjacent-level administrative divisions, and determine a mapping relationship model between adjacent-level administrative divisions according to the hierarchical subordination relationship; Determine a hierarchical association matrix corresponding to the administrative division according to the mapping relationship model; Obtain the original address characters of the administrative division to be standardized, and segment the original address characters to obtain an original address structure corresponding to the original address characters; Obtain the missing complement condition corresponding to the administrative division, determine whether the original address structure satisfies the missing complement condition, and if so, update the original address structure according to the hierarchical association matrix to obtain a standard address structure corresponding to the administrative division to be standardized; Obtain address element sets corresponding to administrative divisions at all levels according to the address element pairs, specifically including: Obtain single-level address element sets corresponding to administrative divisions at all levels according to the address element pairs; For each single-level address element set, determine the number of address elements in the single-level address element set, and determine an index number interval of the administrative division corresponding to the single-level address element set according to the number of address elements; Index the administrative division corresponding to the single-level address element set according to the index number interval to obtain address element sets corresponding to administrative divisions at all levels; Segment the original address characters to obtain an original address structure corresponding to the original address characters, specifically including: Segment the original address characters to obtain an address element sequence composed of multiple address elements corresponding to the original address characters, and sequentially match the multiple address elements with the address element set; When there is a matching address element set, obtain a specified address element in the address element set that matches the address element, and add the specified address element to the original address structure of the original address characters; When there is no matching address element set, obtain an original address structure corresponding to the original address characters according to other address elements in the address element sequence that are after the current address element; Before updating the original address structure according to the hierarchical association matrix, the method further includes: Obtain a target value that satisfies the missing complement condition and a target address element corresponding to the target value; Determine the address element set where the target address element is located, and generate a logical matrix corresponding to the administrative division according to the index number in the address element set; Update the original address structure according to the hierarchical association matrix to obtain a standard address structure corresponding to the administrative division to be standardized, specifically including: Construct a number completion matrix corresponding to the administrative division according to the hierarchical association matrix and the logic matrix; wherein, the number completion matrix is composed of the index numbers of the address element pairs belonging to the administrative division in their corresponding address element sets; Obtain the address element pairs corresponding to each element in the number completion matrix to construct an administrative division completion matrix corresponding to the administrative division; Update the original address structure according to the administrative division completion matrix to obtain the standard address structure corresponding to the administrative division to be standardized.
2. The method for standardizing Chinese address administrative divisions according to claim 1, characterized in that Construct the state space corresponding to the address element set, specifically including: Determine the index numbers corresponding to each administrative division in the address element set, and establish a logic vector equivalent to the index number according to the index number; Construct the state space corresponding to the address element set according to the logic vector.
3. A method for standardizing Chinese address administrative divisions according to claim 1, characterized in that, Determine the mapping relationship model between adjacent-level administrative divisions according to the hierarchical subordination relationship, specifically including: For the state space corresponding to the address element set of each level of administrative division, divide the state space into several state sub-spaces; According to the hierarchical subordination relationship, determine the first mapping relationship between the state sub-space and the state space corresponding to its adjacent-level administrative division, and determine the second mapping relationship between adjacent-level administrative divisions according to the first mapping relationship; Construct a mapping relationship model between adjacent-level administrative divisions according to the second mapping relationship.
4. A method for standardizing Chinese address administrative divisions according to claim 1, characterized in that, Determine the hierarchical association matrix corresponding to the administrative division according to the mapping relationship model, specifically including: Determine the weak hierarchical association matrix corresponding to the administrative division according to the mapping relationship model; According to the weak hierarchical association matrix, determine whether the specified administrative division in the administrative division belongs to the adjacent-level administrative division corresponding to the administrative division; If so, perform iterative calculation on the weak hierarchical association matrix to obtain the hierarchical association matrix corresponding to the administrative division.
5. A Chinese address administrative division standardization system, characterized in that, Applied to a Chinese address administrative division standardization method according to any one of claims 1-4, the system includes: An administrative division element matching module, configured to construct an address element pair corresponding to the administrative division according to the administrative division and the index number corresponding to the administrative division, and obtain an address element set corresponding to each level of administrative division according to the address element pair; Specifically configured to obtain a single-level address element set corresponding to each level of administrative division according to the address element pair; For each single-level address element set, determine the number of address elements in the single-level address element set, and determine the index number interval of the administrative division corresponding to the single-level address element set according to the number of address elements; Index the administrative division corresponding to the single-level address element set according to the index number interval to obtain an address element set corresponding to each level of administrative division; An administrative division hierarchical association module, configured to construct a state space corresponding to the address element set, determine the hierarchical subordination relationship corresponding to adjacent-level administrative divisions, and determine a mapping relationship model between adjacent-level administrative divisions according to the hierarchical subordination relationship; It is also used to determine the hierarchical association matrix corresponding to the administrative division according to the mapping relationship model; An administrative division recognition and conversion module, configured to obtain the original address characters of the administrative division to be standardized, and perform word segmentation on the original address characters to obtain the original address structure corresponding to the original address characters; Specifically, it is used to perform word segmentation on the original address characters to obtain an address element sequence composed of multiple address elements corresponding to the original address characters, and sequentially match the multiple address elements with the address element set; In the case where there is a matching address element set, obtain the specified address element in the address element set that matches the address element, and add the specified address element to the original address structure of the original address characters; In the case where there is no matching address element set, obtain the original address structure corresponding to the original address characters according to other address elements after the current address element in the address element sequence; An administrative division standard complement module, configured to obtain the missing complement condition corresponding to the administrative division, determine whether the original address structure satisfies the missing complement condition, and if so, update the original address structure according to the hierarchical association matrix to obtain the standard address structure corresponding to the administrative division to be standardized; Specifically, before updating the original address structure according to the hierarchical association matrix, the method further includes: Obtain the target value that satisfies the missing complement condition and the target address element corresponding to the target value; Determine the address element set where the target address element is located, and generate the logical matrix corresponding to the administrative division according to the index number in the address element set; Specifically, according to the hierarchical association matrix and the logical matrix, construct the number complement matrix corresponding to the administrative division; wherein, the number complement matrix is composed of the index numbers of the administrative divisions to which the address element pairs belong in their corresponding address element sets; Obtain the address element pairs corresponding to each element in the number complement matrix to construct the administrative division complement matrix corresponding to the administrative division; Update the original address structure according to the administrative division complement matrix to obtain the standard address structure corresponding to the administrative division to be standardized.
6. A Chinese address administrative division standardization device, characterized in that, Applied to a Chinese address administrative division standardization method according to any one of claims 1-4, the device includes: At least one processor; And a memory communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can: Construct the address element pairs corresponding to the administrative division according to the administrative division and the index number corresponding to the administrative division, and obtain the address element sets corresponding to administrative divisions at all levels according to the address element pairs; Construct the state space corresponding to the address element set, determine the hierarchical membership relationship between adjacent-level administrative divisions, and determine the mapping relationship model between adjacent-level administrative divisions according to the hierarchical membership relationship; Determine the hierarchical association matrix corresponding to the administrative division according to the mapping relationship model; Obtain the original address characters of the administrative division to be standardized, and segment the original address characters to obtain the original address structure corresponding to the original address characters; Obtain the missing complement condition corresponding to the administrative division, and determine whether the original address structure satisfies the missing complement condition. If so, update the original address structure according to the hierarchical association matrix to obtain the standard address structure corresponding to the administrative division to be standardized; According to the address element pairs, obtain the address element sets corresponding to administrative divisions at all levels, specifically including: According to the address element pairs, obtain the single-level address element sets corresponding to administrative divisions at all levels respectively; For each single-level address element set, determine the number of address elements in the single-level address element set, and according to the number of address elements, determine the index number interval of the administrative division corresponding to the single-level address element set; According to the index number interval, index the administrative division corresponding to the single-level address element set to obtain the address element sets corresponding to administrative divisions at all levels; Segment the original address characters to obtain the original address structure corresponding to the original address characters, specifically including: Segment the original address characters to obtain an address element sequence composed of multiple address elements corresponding to the original address characters, and sequentially match the multiple address elements with the address element sets; In the case where there is a matching address element set, obtain the specified address element in the address element set that matches the address element, and add the specified address element to the original address structure of the original address characters; In the case where there is no matching address element set, obtain the original address structure corresponding to the original address characters according to the other address elements after the current address element in the address element sequence; Before updating the original address structure according to the hierarchical association matrix, the method further includes: Obtain the target value that satisfies the missing complement condition and the target address element corresponding to the target value; Determine the address element set where the target address element is located, and generate the logical matrix corresponding to the administrative division according to the index number in the address element set; Update the original address structure according to the hierarchical association matrix to obtain the standard address structure corresponding to the administrative division to be standardized, specifically including: Construct the number complement matrix corresponding to the administrative division according to the hierarchical association matrix and the logical matrix; wherein, the number complement matrix is composed of the index numbers of the administrative divisions to which the address element pairs belong in their corresponding address element sets; Obtain the address element pairs corresponding to the elements in the number complement matrix to construct the administrative division complement matrix corresponding to the administrative division; Update the original address structure according to the administrative division complement matrix to obtain the standard address structure corresponding to the administrative division to be standardized.
Citation Information
Patent Citations
Administrative division completion and standardization method for Chinese address
CN111159973A
Province-city-area address information matching method and device, computer equipment and storage medium
CN115185986A