Power industry seven-level address standard library construction method and system
By constructing a seven-level address standard library that aggregates multi-source data, the limitations of data sources and the lack of hierarchical structure in the power industry's address data have been resolved, thereby improving the completeness and accuracy of address data and supporting the refined operation of the power industry.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-03
- Publication Date
- 2026-04-14
AI Technical Summary
The power industry's address data suffers from limitations in data sources, missing hierarchical structures, and poor data quality, making it difficult to provide refined services.
By constructing a seven-level address standard library that aggregates multi-source data, and utilizing data acquisition, preprocessing, hierarchical construction, address integration and improvement layers, combined with confidence quantification of multi-source data and multi-dimensional similarity calculation, the integrity and accuracy of address data are improved.
It has achieved complete coverage and improved accuracy of address data in the power industry, supports refined operations, solved the problems of missing address levels and poor data quality, and improved business service efficiency.
Smart Images

Figure CN121860692A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power industry operation and management, specifically to a method and system for constructing a seven-level address standard library for the power industry. Background Technology
[0002] In the power industry's operation and management, the standardization requirements for address data are becoming increasingly stringent. The accuracy of address data directly impacts key business scenarios such as customer service, equipment management, and emergency repair. However, the current address data system used in the power industry has many problems that urgently need to be addressed. Taking data sources as an example, most existing address databases rely solely on electricity address data within the power marketing system, lacking effective integration and utilization of authoritative external data sources, such as administrative addresses from the Ministry of Civil Affairs and regional division data from the National Bureau of Statistics. This results in a significant shortcoming in address hierarchy coverage, typically only covering the first five levels: "province-city-district / county-street-neighborhood committee." Information at levels crucial for refined services, such as "roads (level six)" and "communities / villages (level seven)," is severely lacking. In actual operations, the impact of this lack of hierarchy is extremely significant. For instance, when conducting centralized meter reading in communities, the lack of level six "road" information makes it difficult for meter readers to plan the optimal reading route, greatly reducing efficiency. In terms of charging pile layout planning, without accurate level seven address data as support, it is impossible to accurately locate potential demand areas, leading to unreasonable resource allocation.
[0003] For example, the existing invention patent application document CN119669491A, entitled "Address Data Governance Method Based on Deep Learning," mentions the use of official data sources such as the Ministry of Civil Affairs and the National Bureau of Statistics. However, its core focus is on achieving standardized governance of address data through deep learning technology. It does not provide a detailed introduction to the multi-source data fusion and construction method for a seven-level address standard library specifically for the power industry. For instance, the data fusion in this patent is limited to general standardization processing and does not establish, for example, a priority rule for power-specific regions based on data from the Ministry of Civil Affairs and supplemented by data from the National Bureau of Statistics. When faced with address conflicts across administrative levels, such as in high-tech zones and economic development zones, it still lacks industry-adaptive solutions.
[0004] In summary, the current power industry address data suffers from problems such as limited data sources, missing hierarchical levels, and poor data quality. There is an urgent need for a method to construct a seven-level address standard library based on multi-source data aggregation to improve the integrity, accuracy, and usability of address data and support the refined operation of the power industry.
[0005] In summary, existing power industry address data suffers from technical problems such as limited data sources, missing hierarchical levels, and poor data quality. Summary of the Invention
[0006] The technical problem to be solved by this invention is: how to solve the technical problems of limited data sources, missing levels, and poor data quality in the existing power industry address data.
[0007] This invention solves the above-mentioned technical problems by employing the following technical solution: A method for constructing a seven-level address standard library for the power industry includes:
[0008] S1. Using the data acquisition layer, obtain the top five levels of administrative address data and the associated data of the five levels of administrative addresses, and process them to obtain multi-source address data, including: codes, government platform address data, electricity address data, and map vendor address data; S2. Using the data preprocessing layer, address segmentation, NLP, and regular expressions are employed to clean, reduce noise, and perform structure transformation operations on multi-source address data to obtain preprocessed multi-source address data. S3. By utilizing the hierarchical construction layer and integrating government platform and map merchant data, based on preprocessed multi-source address data, and supplementing with level six roads and level seven community addresses, a seven-level address framework is constructed. S4. Using the address integration layer, establish a priority comparison model for the top five levels of addresses, obtain and integrate data from the Ministry of Civil Affairs and the National Bureau of Statistics to form basic data for the top five levels of standard addresses. S5. Utilize the address improvement layer, integrate government platform data, and perform initial construction of the first seven-level address database based on the seven-level address framework. Through electricity address data and preprocessed multi-source address data, perform model mutual verification to obtain the improved sixth and seventh level addresses. S6. Utilize the standard library output layer to output the power industry's level 7 standard address library based on the improved level 6 and level 7 addresses.
[0009] This invention constructs a seven-level address standard library for the power industry based on multi-source data aggregation. The aim is to improve the accuracy and consistency of address data in the power industry by integrating multi-source data, providing efficient and accurate address service support for power marketing, production, and other business systems. This invention can improve the integrity, accuracy, and usability of address data, supporting the refined operation of the power industry.
[0010] In a more specific technical solution, in S2, non-hierarchical information is removed and the description is unified according to the hierarchical name extraction rules, thus completing the preprocessing operation of the electricity address data; The confidence weights of each data source at different levels and in different regions are quantified, the confidence of each data source is quantified, and a dimension scoring-weight calculation evaluation model is constructed. In the process of quantitatively assessing the confidence level of each data source, single-dimensional score statistics are performed, and scores for each dimension are calculated through field sampling. A weight calculation model is constructed, and the entropy weight method is used to determine the weight of each dimension. The comprehensive confidence weight of the data source is calculated by combining the single-dimensional scores.
[0011] This invention achieves complete coverage of seven levels of addresses, meeting the needs of refined power operation. It overcomes the limitations of existing technologies that rely on a single data source, constructing a multi-source data system comprising the Ministry of Civil Affairs, the National Bureau of Statistics, government service platforms, map providers, and electricity usage addresses. Data sources are prioritized according to power business needs: data from the Ministry of Civil Affairs serves as the first five levels of administrative benchmarks (ensuring administrative authority); data from the National Bureau of Statistics supplements special cross-level areas such as "High-tech Zones / Economic Development Zones"; data from the government service platform fills in gaps in village and group addresses; and data from map providers and the State Grid GIS platform completes the sixth level of road information. Compared to the "general standardized processing" of patent CN119669491A, this invention achieves "precise matching of data sources with power business scenarios." For example, by using the addresses of High-tech Zones supplemented by the National Bureau of Statistics, it solves the problem of "no corresponding administrative code for High-tech Zone users opening accounts" in electricity usage addresses.
[0012] In a more specific technical solution, S4, the calculation of the comprehensive confidence weight of the data source is performed. W During the process, the entropy values of each dimension are calculated. This reflects the contribution of dimensional information: , Among them, the number of data sources , For the first The data source in the first The percentage of scores for each dimension; Based on the entropy values of each dimension Calculate dimensional weights : ; Calculate the overall confidence weight of the data source :
[0013] In the formula, For the first The first data source Dimensional standardized score.
[0014] In a more specific technical solution, S4 integrates the basic data of the top five standard addresses based on the priority comparison model; aligns the fields of the data from the Ministry of Civil Affairs and the National Bureau of Statistics to obtain aligned data; and performs data coverage comparison and data difference comparison.
[0015] In a more specific technical solution, in S4, when both the Ministry of Civil Affairs and the National Bureau of Statistics have address data for the same administrative unit, the final data is determined based on preset conditions during the difference comparison operation; NLP semantic similarity calculation is used to detect name differences; and encoding difference processing is performed.
[0016] In a more specific technical solution, S5 performs multi-source mutual verification based on name comparison. Taking the initial version of the seven-level address database as a benchmark, it completes the model mutual verification operation by comparing the consistency of names at the same level and verifying the rationality of the superior-subordinate relationship. It also calculates the consistency score by combining the confidence weight of the data source, identifies and corrects the problems in the initial version of the seven-level address database.
[0017] This invention utilizes a multi-source mutual verification model to resolve data conflicts and improve address credibility. Addressing address differences from different data sources, this invention innovatively designs a mutual verification mechanism combining "quantitative confidence assessment + multi-dimensional similarity calculation": It uses entropy weighting to determine the weights of each data source in different regions (urban / rural) and at different levels, and combines algorithms such as edit distance and Jaccard similarity to calculate consistency scores. Ultimately, this achieves accurate identification and correction of conflicting data, effectively avoiding problems such as incorrect electricity bill collection and inaccurate repair addresses caused by address conflicts.
[0018] In a more specific technical solution, during the process of comparing name consistency at the same level, name similarity is calculated by fusing edit distance similarity and Jaccard similarity, covering character differences and semantic associations; The edit distance similarity is calculated using the following logic:
[0019] In the formula, This is the name of the initial version of the library to be verified. To supplement the names of data sources at the same level, To edit distance, The length of the name in characters; Calculate the Jaccard similarity using the following logic: ; Using the following logic, calculate the fusion similarity and take the average of the two as the final name similarity at the same level: .
[0020] In a more specific technical solution, during the calculation of the confidence-weighted consistency score, the overall consistency score of names at the same level is calculated by combining the confidence weights of each data source. :
[0021] In the formula, For the first The similarity between each data source and the initial version of the library at the same level of integration. Assign confidence weights to the corresponding data sources; During the process of verifying the rationality of hierarchical relationships, hierarchical relationship representation and matching are performed; relationship consistency score is calculated. Using the following logic, combined with the confidence weight of the data source, calculate the relationship consistency score. :
[0022] In the formula, It is an indicator variable.
[0023] In a more specific technical solution, S5 involves address database omissions based on high-confidence names. This is achieved by comparing names from multiple sources with those from the initial seven-level address database to identify addresses not covered in the initial version, and then supplementing them based on names from high-confidence data sources. Specifically, a reverse matching operation is used, comparing name combinations from multiple sources with the seven-level name combinations from the initial seven-level address database to identify missing addresses. For missing addresses, the hierarchical relationships are completed. Finally, missing address encoding and assignment operations are performed, and the supplementation effect is evaluated.
[0024] This invention features a hierarchical design tailored to the power industry, breaking through the limitations of a five-level system. For the first time, this invention constructs a seven-level address system for the power industry: "Province-City-District / County-Street-Neighborhood Committee-Road-Community / Village." The sixth level, "Road," is adapted for fault repair route planning (repair personnel can quickly locate fault areas using road information), while the seventh level, "Community / Village," is adapted for scenarios such as centralized meter reading (meter readers plan routes by community group) and charging pile deployment (precise site selection based on community electricity demand). Compared to existing technologies that only cover the first five levels of addresses, this invention significantly enhances the adaptability to power business scenarios, providing stronger data support for the development of power business services.
[0025] In a more specific technical solution, a system for constructing a seven-level address standard library for the power industry includes: The data acquisition layer is used to acquire the top five levels of administrative address data and the associated data of the five levels of administrative addresses, and processes them to obtain multi-source address data, including: codes, government platform address data, electricity address data, and map vendor address data; The data preprocessing layer is used to perform cleaning, noise reduction, and structure transformation operations on multi-source address data using address segmentation, NLP, and regular expressions to obtain preprocessed multi-source address data. The data preprocessing layer is connected to the data acquisition layer. The hierarchical construction layer is used to integrate government platform and map merchant data. Based on preprocessed multi-source address data, it supplements the sixth-level road and seventh-level community addresses to construct a seven-level address framework. The hierarchical construction layer is connected to the data preprocessing layer. The address integration layer is used to establish a priority comparison model for the top five levels of addresses, acquire and integrate data from the Ministry of Civil Affairs and the National Bureau of Statistics to form the basic data for the top five levels of standard addresses. The address integration layer is connected to the hierarchical construction layer. The address improvement layer is used to integrate government platform data. Based on the seven-level address framework, it performs initial construction operations on the first seven levels of the address database. Through electricity address data and pre-processed multi-source address data, it performs model mutual verification operations to obtain improved level six and level seven addresses. The address improvement layer is connected to the address integration layer. The standard library output layer is used to output a power industry level 7 standard address library based on the improved level 6 and level 7 addresses. The standard library output layer is connected to the address improvement layer.
[0026] The present invention has the following advantages over the prior art: This invention constructs a seven-level address standard library for the power industry based on multi-source data aggregation. The aim is to improve the accuracy and consistency of address data in the power industry by integrating multi-source data, providing efficient and accurate address service support for power marketing, production, and other business systems. This invention can improve the integrity, accuracy, and usability of address data, supporting the refined operation of the power industry.
[0027] This invention achieves complete coverage of seven levels of addresses, meeting the needs of refined power operation. It overcomes the limitations of existing technologies that rely on a single data source, constructing a multi-source data system comprising the Ministry of Civil Affairs, the National Bureau of Statistics, government service platforms, map providers, and electricity usage addresses. Data sources are prioritized according to power business needs: data from the Ministry of Civil Affairs serves as the first five levels of administrative benchmarks (ensuring administrative authority); data from the National Bureau of Statistics supplements special cross-level areas such as "High-tech Zones / Economic Development Zones"; data from the government service platform fills in gaps in village and group addresses; and data from map providers and the State Grid GIS platform completes the sixth level of road information. Compared to the "general standardized processing" of patent CN119669491A, this invention achieves "precise matching of data sources with power business scenarios." For example, by using the addresses of High-tech Zones supplemented by the National Bureau of Statistics, it solves the problem of "no corresponding administrative code for High-tech Zone users opening accounts" in electricity usage addresses.
[0028] This invention utilizes a multi-source mutual verification model to resolve data conflicts and improve address credibility. Addressing address differences from different data sources, this invention innovatively designs a mutual verification mechanism combining "quantitative confidence assessment + multi-dimensional similarity calculation": It uses entropy weighting to determine the weights of each data source in different regions (urban / rural) and at different levels, and combines algorithms such as edit distance and Jaccard similarity to calculate consistency scores. Ultimately, this achieves accurate identification and correction of conflicting data, effectively avoiding problems such as incorrect electricity bill collection and inaccurate repair addresses caused by address conflicts.
[0029] This invention features a hierarchical design tailored to the power industry, breaking through the limitations of a five-level system. For the first time, this invention constructs a seven-level address system for the power industry: "Province-City-District / County-Street-Neighborhood Committee-Road-Community / Village." The sixth level, "Road," is adapted for fault repair route planning (repair personnel can quickly locate fault areas using road information), while the seventh level, "Community / Village," is adapted for scenarios such as centralized meter reading (meter readers plan routes by community group) and charging pile deployment (precise site selection based on community electricity demand). Compared to existing technologies that only cover the first five levels of addresses, this invention significantly enhances the adaptability to power business scenarios, providing stronger data support for the development of power business services.
[0030] This invention solves the technical problems of limited data sources, missing hierarchical structures, and poor data quality in existing power industry address data. Attached Figure Description
[0031] Figure 1 This is a schematic diagram illustrating the basic steps of a method for constructing a seven-level address standard library for the power industry, as described in Embodiment 1 of the present invention. Detailed Implementation
[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0033] Example 1 like Figure 1 As shown, the present invention provides a method for constructing a seven-level address standard library for the power industry, which includes the following basic steps: S1. Based on the priority comparison model, integrate the address data of the first five levels; In this embodiment, the overall architecture of the present invention includes, but is not limited to: a data acquisition layer, a data preprocessing layer, a hierarchy construction layer, an address integration layer, an address improvement layer, and a standard library output layer; Specifically, the data acquisition layer is used to obtain the top five levels of administrative address data and codes from the Ministry of Civil Affairs and the National Bureau of Statistics, as well as government platform address data, electricity address data, and map merchant address data. By utilizing a data preprocessing layer, multi-source address data is cleaned, denoised, and structurally transformed using techniques such as address segmentation, NLP, and regular expressions. By utilizing a hierarchical structure and integrating government platforms and map vendor data, a six-level (road) and a seven-level (community) address framework are initially constructed based on the five-level framework. Using the address integration layer, a top five address priority comparison model is established, integrating data from the Ministry of Civil Affairs and the National Bureau of Statistics to form a standard address base for the top five levels; By utilizing the address improvement layer and integrating data from the government affairs platform, a preliminary seven-level address database is constructed. Then, by using electricity address data and a multi-source data mutual verification model, the sixth and seventh level addresses are improved. Using the standard library output layer, a seven-level standard address library for the power industry is output, which includes “province (level 1) - city (level 2) - district / county (level 3) - street (level 4) - neighborhood committee (level 5) - road (level 6) - community / village (level 7)”.
[0034] In the integration of the first five levels of address data in this embodiment, data acquisition and field alignment are performed; Specifically, obtain the address data of the top five levels of administrative divisions from the Ministry of Civil Affairs and the National Bureau of Statistics. The data fields include, but are not limited to: province code, province name, city code, city name, district code, district name, street code, street name, neighborhood committee code, and neighborhood committee name. Field alignment was performed on the two types of data, using the data fields from the Ministry of Civil Affairs as the benchmark, to ensure that the administrative level of the two types of data corresponds one-to-one with the basic fields, laying the foundation for subsequent comparison.
[0035] In the construction and execution of the priority comparison model in this embodiment, a priority comparison model is constructed based on data from the Ministry of Civil Affairs and supplemented by data from the National Bureau of Statistics. The specific comparison logic and calculation process are as follows: (1) Data coverage comparison: The system iterates through the special fields in the statistics bureau's data, filtering out records not included in the first five levels of data from the Ministry of Civil Affairs (e.g., "XX City, XX District / County, XX High-tech Zone"). Using the "Main Level of Special Area" field, it locates the corresponding main level of the Ministry of Civil Affairs (e.g., "XX City, XX District / County"). This special area is then added to the first five levels of data as an "Extended Unit under the District / County Level," using the relevant codes for the corresponding statistical area level. The data source is marked as "Supplemented by the Statistics Bureau," and associated with the main level neighborhood committee (e.g., neighborhood committees within the High-tech Zone still use the Ministry of Civil Affairs code). This ensures that subsequent sixth-level (road) and seventh-level (community) addresses can be correctly connected.
[0036] (2) Data difference comparison: When both the Ministry of Civil Affairs and the National Bureau of Statistics have address data for the same administrative unit (such as the same neighborhood committee), they need to compare the differences between the two (such as differences in name description or coding), and determine the final data based on the principle that "the Ministry of Civil Affairs has higher authority." The specific steps are as follows: Name difference detection: NLP semantic similarity calculation is adopted. First, the core fields (such as "XX Street XX Neighborhood Committee") are extracted by Jieba word segmentation. Then, the core fields are converted into 100-dimensional vectors by Word2Vec model and cosine similarity is calculated.
[0037] formula:
[0038] If the similarity is ≥0.95, it is judged as "identical in name description," and the Ministry of Civil Affairs name is retained. If the similarity is <0.95, a dedicated verification for the electricity sector is initiated—the historical electricity address records of electricity users in the area are queried. If more than 80% of the user addresses match the Ministry of Civil Affairs name, the Ministry of Civil Affairs data is retained, and the Statistics Bureau data is recorded as a note. Otherwise, verification is conducted in conjunction with the community office's telephone verification, and the final verification result shall prevail. Handling of coding differences: If the Ministry of Civil Affairs code and the National Bureau of Statistics code of the same administrative unit are inconsistent, the Ministry of Civil Affairs code shall be used as the standard code and the National Bureau of Statistics code shall be retained as a "historical code" field to facilitate subsequent data traceability.
[0039] Through the above comparison model, the first five levels of standard address data are finally formed, which are "completely covered and authoritatively unified", providing a foundation for the construction of the subsequent seven-level address database.
[0040] S2. Conduct preliminary construction of the first seven levels of standard address database; In this embodiment, based on the integration of authoritative data from the Ministry of Civil Affairs and the National Bureau of Statistics to complete the construction of the first five levels of standard administrative addresses (province-city-district-street-neighborhood committee, including administrative standard code values), and by integrating address data from the government affairs platform, the sixth level of road addresses and the seventh level of community / village addresses are supplemented, standardized, and hierarchically linked, forming an initial version of the seventh-level standard address database. This primarily addresses issues related to cross-data source address matching, information gaps, and name standardization. Specific operations include, but are not limited to: First five levels of address normalization mapping; In this embodiment, to achieve accurate association between the first five levels of non-standard addresses and the standard library, a two-level mapping algorithm of "semantic similarity calculation + spatial verification" is designed, and the specific operation is as follows: In this embodiment, address feature vectors are constructed. Specifically, five levels of features—province, city, district / county, street, and neighborhood committee—are extracted from the government platform address (denoted as A) and the standard address (denoted as S) to construct a feature vector with a dimension of 5. The vector value is calculated using the "synonym weighting method": the feature vector value at a certain level is equal to the sum of the matching weights of all synonyms of that level's address with the standard name. The weight is 1.0 for a complete match and 0.6-0.9 for a partial match, dynamically adjusted according to the keyword overlap.
[0041] In this embodiment, a comprehensive similarity calculation is performed; specifically, semantic similarity is calculated, and the overall matching degree of the five-level addresses is calculated using weighted cosine similarity, with the following formula:
[0042] in, The hierarchical weighting coefficients are defined as follows: 0.3 for streets and neighborhood committees, 0.2 for districts and counties, and 0.1 for provinces and cities. , These are the addresses of the government platform and the standard address, respectively. i Level eigenvector values.
[0043] The spatial verification coefficient is calculated based on GIS coordinates to determine the spatial distance attenuation coefficient of geographic entities. The formula is as follows:
[0044] in, d The straight-line distance (unit: km) between the center point of the geographic entity corresponding to the government platform address and the standard address. This is the distance attenuation factor (0.5 for urban areas and 0.2 for rural areas).
[0045] The overall matching degree is calculated by fusing semantic similarity and spatial verification coefficient, using the following formula:
[0046] In this embodiment, mapping decision rules are set; Specifically, based on the overall matching degree Retrieve the value and execute the three-level mapping strategy: High matching The system automatically associates the government affairs platform address with the corresponding standard address code; Matching Triggering a manual review process and synchronously recording mapping rules for algorithm optimization; low matching Marked as "address to be supplemented", the detailed information will be retrieved through the government affairs platform interface and then rematched.
[0047] Practical verification shows that the accuracy of this mapping method reaches 95.8%, which is 27.6% higher than that of traditional string matching methods.
[0048] In this embodiment, a six-level road address completion is performed; specifically, a "multi-source retrieval + weighted fusion" strategy is adopted to complete road information and link it with the addresses of the upper-level fifth-level neighborhood committees and the lower-level seventh-level communities / villages.
[0049] In this embodiment, multi-source road data retrieval is performed; Map search: Call the POI search service, take the center point coordinates of the community as the center and 1km as the search radius, and obtain information such as the name, grade, start and end points and center point coordinates of the surrounding roads; State Grid GIS Platform Search: Call the internal address association interface, enter the community name, and return the road information around the community with the center point coordinates of the community as the center and a search radius of 1km.
[0050] In this embodiment, the road-community / village address association algorithm is executed; Specifically, the correlation between roads and the seventh-level address communities / villages is calculated using a frequency-distance weighted method, with the following formula:
[0051] in, For roads Frequency percentage in multi-source search results For roads The shortest straight-line distance to cell c with a level 7 address. (Maximum search distance).
[0052] Roads with a correlation score of ≥0.6 are selected as the candidate set of Level 6 addresses. After sorting by road level (main road > secondary road > branch road > rural road), the TOP1 is taken as the main road corresponding to the community, and the rest are taken as related roads, so as to achieve accurate connection between Level 6 roads and Level 7 community / village addresses.
[0053] In this embodiment, a six- or seven-level address encoding rule is established; Specifically, since the address data of the seventh-level residential communities / villages on the government affairs platform and the supplementary sixth-level road address data lack standard codes, the coding rules for sixth-level roads and seventh-level residential communities / villages are formulated as follows, taking into account power industry standards and business needs: Level 6 road address code (15 digits) Encoding structure: It adopts an "8-digit five-level association code + 7-digit road-specific code" concatenation method, where the definition of each segment completely corresponds to the code value allocation standard in the attachment: (1) 8-bit five-level association code Value source: The last 8 digits are extracted from the first five levels of the 12-digit standard administrative code (corresponding to the 5th to 12th digits of the code), which are accurately associated with the three-level administrative units of "district / county-street-neighborhood committee". The sixth-level code prefix must be directly mapped to the last 8 digits of the fifth-level neighborhood committee code to ensure the traceability of administrative affiliation.
[0054] Example verification: The first five levels of code "340104002011" (Meishan Road Community, Sanli'an Street, Shushan District, Hefei City, Anhui Province) are used as the prefix for all roads in the community. The last 8 digits are extracted as "04002011" as required in the attachment.
[0055] (2) 7-digit road-specific code It consists of a 2-digit city code + a 5-digit sequence code: The two-digit city code adopts the administrative city numerical code of the Ministry of Civil Affairs and the National Bureau of Statistics. Taking Anhui as an example, it is shown in the attached table below:
[0056] 5-digit sequence code generation rules: Following the principle of "unique within the city / prefecture," roads are prioritized according to the order of "main roads → secondary roads → branch roads → rural roads," using a sequence code of 00001-99999. For example, Huangshan Road, Honggang Road, and Jinzhai Road in Hefei City, and Guyang Road in Bengbu City are coded as 0100001, 0100002, 0100003, and 0300001, respectively. The complete road code is obtained by concatenating the first 8 digits, as shown in the example below: 040030040100001 indicates: Huangshan Road, Huangshan Road Community, Daoxiangcun Street, Shushan District, Hefei City; 040030040100002 indicates: Honggang Road, Huangshan Road Community, Daoxiangcun Street, Shushan District, Hefei City; 110110030100003 indicates: Jinzhai Road, Weigang Community, Tongan Street, Baohe District, Hefei City.
[0057] Level 7 Community / Village Address Code (16 digits) The coding structure adopts a "10-digit five-level association code + 6-digit community / village group address exclusive code" structure to achieve accurate association with neighborhood committees and electricity customers.
[0058] (1) Method for obtaining the 10-digit five-level association code: Remove the first two provincial codes from the first five-level 12-digit code (retain the 3rd to 12th digits) to form a four-level association code containing "city-district-street-neighborhood committee". The seven-level code prefix must cover the municipal administrative unit to ensure the uniqueness of cross-district and county addresses.
[0059] Example comparison: After removing the provincial code "34" from the first five levels of the code "340104002011", the 10-digit association code is "0104002011", which corresponds to "Hefei City - Shushan District - Sanli'an Street - Meishan Road Community".
[0060] (2) The 6-digit Level 7 community / village group address exclusive code is the unique identifier of the Level 7 community / village group address under the same neighborhood committee. Within the same neighborhood committee, it is prioritized according to the recommended "construction time priority" principle. Supplement the power business adaptation rules, and at the database level, create a unique index of "10-digit association code + 6-digit exclusive code". The construction example is as follows: The address of Pacific Senghuo Mansion on Xiangzhang Avenue, Jinxiu Yiyuan Community Residents Committee, Taohua Town, Feixi County, Hefei City is 0123102012000001. The address of Dingyuan Mansion on Mingzhu Road, Jinxiu Yiyuan Community Residents Committee, Taohua Town, Feixi County, Hefei City is 0123102012000002. The address of Jin'an Garden, Huangshan Road, Nanyuan Community Residents Committee, Wuhu Road Subdistrict, Baohe District, Hefei City is 0111003006000001.
[0061] Initial version of the seven-level standard address database integration; Through the multi-source data fusion and standardization processing of the above steps, the initial version of the seven-level standard address library achieves complete connection of seven levels of addresses: "province-city-district / county-street-neighborhood committee-road-community / village". This initial version of the library provides a unified and standard address benchmark for various business systems in the power industry, which can directly support business scenarios such as customer address verification and management, facility location, etc., and at the same time lay the foundation for subsequent iterative optimization of the address library.
[0062] S3, a comprehensive seven-level address database with multi-source mutual verification; In this embodiment, multi-source data preprocessing is performed; Preprocessing of electricity address data is performed; specifically, the electricity addresses exist in long text format (e.g., "Unit 2, Building 1, Jin'an Garden, Huangshan Road, Nanyuan Community Residents Committee, Wuhu Road Subdistrict, Baohe District, Hefei City"). The core is to extract the seven-level names and remove non-hierarchical information such as building and unit names, which is achieved using "regular expressions + keyword matching". Hierarchical name extraction rules: Based on the characteristics of address names in the power industry, define keywords and extraction templates for each level to ensure accurate segmentation; Non-hierarchical information removal: Filter redundant content such as "building number, unit number, and door number" using keywords. For example, the regular expression rule is: (?<=[community / village / group])(. ?[Building Unit Number\d]+), remove “Building 1, Unit 2” from the example; Consistent wording: Standardize different wording at the same level (e.g., “Wuhu Road Street” → “Wuhu Road Subdistrict”, “Nanyuan Community” → “Nanyuan Neighborhood Committee”) to ensure consistency with the initial version of the database.
[0063] Map providers and the State Grid GIS platform perform address data preprocessing; in this embodiment, map providers return structured hierarchical names (e.g., "Province: Anhui Province, City: Hefei City, District / County: Baohe District, Street: Wuhu Road Street, Community: Nanyuan Community Residents Committee, Road: Huangshan Road, POI Name: Jin'an Garden"). Preprocessing only needs to complete "representation alignment + name completion". Alignment of descriptions: Match the hierarchical names with the keywords in the initial version of the database, such as "community" → "neighborhood committee" in the initial version of the database ("Nanyuan Community" → "Nanyuan Community Residents' Neighborhood Committee"), and "village group" in the State Grid GIS platform → "XX Village XX Group" ("Nanyuan Group" → "Nanyuan Village 01 Group"). Name completion: If a name at a certain level is missing, it will be completed based on the name of the parent level to ensure that the names at all seven levels are complete.
[0064] Quantitative evaluation of confidence of multiple data sources; In this embodiment, since the accuracy of the names and update frequency of the three types of supplementary data sources are different, it is necessary to first quantify the confidence weight of each data source in "different levels and different regions" to provide priority basis for subsequent mutual verification and construct an evaluation model of "dimension scoring-weight calculation".
[0065] In this embodiment, the confidence level assessment dimensions are designed based on the electricity service's requirements for address names (e.g., the community name must be exactly the same as the customer's electricity address, and the road name must be suitable for emergency repair routes). Four core assessment dimensions are defined:
[0066] In this embodiment, the confidence weight of the data source is calculated; Single-dimensional score statistics: Scores for each dimension are calculated through on-site sampling (1,000,000 samples from each data source).
[0067] Weighting calculation model: The "entropy weight method" is used to determine the weight of each dimension, and then the overall confidence weight of the data source is calculated by combining the single-dimensional scores (denoted as...). The formula is: Calculate the entropy value of each dimension This reflects the contribution of dimensional information: ,in (Number of data sources) For the first The data source in the first The percentage of scores for each dimension; Calculate dimensional weights :
[0068] The results are as follows: accuracy 0.35, update timeliness 0.25, power business adaptability 0.30, regional coverage advantage 0.10; Calculate the overall confidence weight of the data source : ( For the first The first data source Dimensional standardized score (value 0-1), the weighted result is: electricity address data, city / region weight. Rural area weight Suitable for verifying the names of urban residential communities; map data, city area weight. Rural area weight Suitable for verifying urban road names; data from the State Grid GIS platform, with urban area weights. Rural area weight It is applicable to the verification of village and group names in rural areas.
[0069] In this embodiment, multi-source mutual verification based on name comparison is implemented. Specifically, taking the initial version of the seven-level address database as a benchmark, the consistency score is calculated by combining the data source confidence weight with two mutual verification dimensions: "consistency comparison of names at the same level" and "reasonableness verification of hierarchical relationship". Problems in the initial version of the database are identified and corrected.
[0070] To address the issue of non-standard names, we conduct mutual verification of names at the same level. Specifically, for abbreviations in the initial version of the database's level 7 names (such as "Taipingsen" → "Pacific Forest Mansion"), we compare names at the same level from multiple data sources to select standard names. The core of this approach is "similarity calculation + confidence weighting".
[0071] Perform name similarity calculation; specifically, use "edit distance similarity + ... Jaccard Similarity calculations are performed in a fusion manner, covering character differences and semantic associations: Edit distance similarity (a measure of the cost of character modification):
[0072] in, This is the name of the initial version of the library to be verified. To supplement the names of data sources at the same level, This is the edit distance (the minimum number of insertions, deletions, and replacements required). The length of the name in characters; Example: .
[0073] Jaccard similarity (a measure of keyword overlap):
[0074] The "word segmentation" uses Jieba (which loads a dictionary specifically for electricity addresses), for example: , .
[0075] Fusion Similarity: The average of the two values is taken as the final name similarity at the same level. The formula is:
[0076] In the example (Initial assessment indicates a mismatch; further analysis using other data sources is required.)
[0077] In this embodiment, the confidence-weighted consistency score is used. By combining the confidence weights of each data source, a comprehensive consistency score (denoted as ) is calculated for names at the same level. The formula is:
[0078] in, For the first The similarity between each data source and the initial version of the library at the same level of integration. Assign confidence weights to the corresponding data sources; Example: Verification of the urban area "Taipingsen": Electricity address data: The contribution value is approximately 0.06 (0.38 × 0.17). Map data: , ; Data from the State Grid GIS platform: The contribution value is approximately 0.04 (0.27 × 0.16). Overall consistency score (Low score, the initial version name is deemed non-standard and needs to be corrected).
[0079] (3) Nomenclature standardization correction rules according to Classify correction levels and select standard names based on multiple source names:
[0080] The example uses "Taipingsen". The system automatically replaced the name with "Pacific Forest Mansion" which was consistent across multiple data sources, increasing the name standardization rate from 96.8% in the initial version to 99.2% after the correction.
[0081] In this embodiment, the legitimacy of the superior-subordinate relationship is mutually verified, thus resolving the issue of deviation in the subordinate relationship. The initial version of the database may contain errors such as "the community belongs to the wrong neighborhood committee" (e.g., "Jin'an Garden" belongs to "Xiyuan Community" but should actually belong to "Meishan Road Community"). The rationality of the affiliation can be verified by comparing the "superior name - subordinate name" relationship of multiple data sources.
[0082] (1) Representation and matching of hierarchical relationships The hierarchical relationship is represented as a pair of "(superior level name, subordinate level name)", such as "(Meishan Road Community, Jin'an Garden)". The core of mutual verification is to compare the consistency of this relationship pair in multi-source data.
[0083] (2) Calculation of Relationship Consistency Score By combining the confidence weights of the data source, a relationship consistency score is calculated (denoted as ). The formula is:
[0084] in, As an indicator variable: if the hierarchical relationship of the i-th data source is consistent with the initial version of the library, ,otherwise ; Example (Initial relationship "(Xiyuan Community, Jin'an Garden)", city area): Electricity usage address data: Relationship "(Meishan Road Community, Jin'an Garden)" ; Map data: Relationship "(Meishan Road Community, Jin'an Garden)" ; Data from the State Grid GIS platform: Relationship "(Meishan Road Community, Jin'an Garden)" ; Relationship Consistency Score (Incorrect determination of affiliation).
[0085] (3) Rules for modifying affiliation according to Determine the correction strategy to ensure a reasonable superior-subordinate relationship:
[0086] In the example The system automatically corrected the affiliation relationship to be consistent across multiple sources: “(Meishan Road Community, Jin'an Garden)”. After the correction, the accuracy of the affiliation relationship improved from 96.8% in the initial version to 99.3%.
[0087] In this embodiment, missing information in the address database based on high-confidence names is supplemented. Specifically, by comparing the names of multi-source data with those of the initial database, addresses not covered by the initial database (such as newly built residential areas and remote villages) are identified, and the database is supplemented based on names from high-confidence data sources to ensure the integrity of the address database.
[0088] In this embodiment, missing address identification is performed; Using the "reverse matching method," based on the initial database's "seven-level name combination" (province + city + district / county + street + neighborhood committee + road + residential area / village), the name combinations from multiple sources are compared. The steps are as follows: Construct the initial version of the database with a "seven-level name combination index": use "province-city-district-street-neighborhood committee-road-community / village group" as the unique index key to build a hash table (query efficiency O(1)); Multi-source data matching: Matching the "seven-level name combination" of the three types of preprocessed data sources with the index key; Missing address determination: If a name combination has no match in the initial version of the database index, but exists in ≥2 types of data sources (ensuring it is not an isolated case), it is marked as a "missing address"; Example: Both the map provider and the State Grid GIS platform contain the name "Huangshan Road Jinxiuyuan, Meishan Road Community, Sanli'an Street, Shushan District, Hefei City". The initial version of the database does not contain this name combination, so it is determined to be missing.
[0089] In this embodiment, missing address hierarchical relationships are filled in; For any missing addresses identified, the complete seven-level name combination and hierarchical relationship must be completed, following the rules: Name combination completion: If a name at a certain level of the address is missing, prioritize using "State Grid GIS Platform Name" (rural area) or "Electricity Address Data Name" (urban area) to complete it; Determining hierarchical relationships: Based on the relationship pairs from high-confidence data sources, determine the affiliation of missing addresses (e.g., if "Jinxiuyuan" belongs to "Meishan Road Community" in the electricity address data, then complete the relationship "(Meishan Road Community, Jinxiuyuan)").
[0090] In this embodiment, missing address encoding is performed; According to the encoding rules mentioned above (Level 6 15 bits: 8-bit Level 5 association code + 7-bit road-specific code; Level 7 16 bits: 10-bit Level 5 association code + 6-bit road-specific code), assign values to the missing addresses: Level 6 road coding: If the address is missing and contains a road name, it is generated by "8-digit level 5 association code (taken from the last 8 digits of the initial level 5 code corresponding to the neighborhood committee name) + 7-digit road-specific code (city code + sequence code)"; The seventh-level address code is generated by "10-digit fifth-level association code (taken from the first version of the fifth-level code corresponding to the neighborhood committee name, after removing the provincial code and the last 10 digits) + 6-digit exclusive code (maximum sequence under the same neighborhood committee + 1)"; Example: If the address "Huangshan Road Jinxiuyuan, Meishan Road Community, Sanli'an Street, Shushan District, Hefei City" is missing, and the level 5 code is "340104002011", then the level 7 code is "0104002011000003" (the current maximum sequence under the same neighborhood committee is 000002).
[0091] In this embodiment, the supplementary effect is evaluated; Specifically, the supplementary effect is quantified through "regional coverage rate," using the following formula:
[0092] The results are as follows:
[0093] In this embodiment, the quality of the improved seven-level standard address database is verified. Specifically, based on three core indicators—name standardization, relationship rationality, and completeness—and combined with the power business scenario, the quality is verified, and the results are as follows:
[0094] Verification results show that the improved Level 7 standard address database fully meets the power industry's requirements for address services that are "accurate in name, clear in relationship, and comprehensive in coverage," and can directly support the efficient operation of core businesses such as marketing meter reading, customer location, and fault repair.
[0095] In summary, this invention constructs a seven-level address standard library for the power industry based on multi-source data aggregation. The aim is to improve the accuracy and consistency of address data in the power industry by integrating multi-source data, providing efficient and accurate address service support for power marketing, production, and other business systems. This invention can improve the integrity, accuracy, and usability of address data, supporting the refined operation of the power industry.
[0096] This invention achieves complete coverage of seven levels of addresses, meeting the needs of refined power operation. It overcomes the limitations of existing technologies that rely on a single data source, constructing a multi-source data system comprising the Ministry of Civil Affairs, the National Bureau of Statistics, government service platforms, map providers, and electricity usage addresses. Data sources are prioritized according to power business needs: data from the Ministry of Civil Affairs serves as the first five levels of administrative benchmarks (ensuring administrative authority); data from the National Bureau of Statistics supplements special cross-level areas such as "High-tech Zones / Economic Development Zones"; data from the government service platform fills in gaps in village and group addresses; and data from map providers and the State Grid GIS platform completes the sixth level of road information. Compared to the "general standardized processing" of patent CN119669491A, this invention achieves "precise matching of data sources with power business scenarios." For example, by using the addresses of High-tech Zones supplemented by the National Bureau of Statistics, it solves the problem of "no corresponding administrative code for High-tech Zone users opening accounts" in electricity usage addresses.
[0097] This invention utilizes a multi-source mutual verification model to resolve data conflicts and improve address credibility. Addressing address differences from different data sources, this invention innovatively designs a mutual verification mechanism combining "quantitative confidence assessment + multi-dimensional similarity calculation": It uses entropy weighting to determine the weights of each data source in different regions (urban / rural) and at different levels, and combines algorithms such as edit distance and Jaccard similarity to calculate consistency scores. Ultimately, this achieves accurate identification and correction of conflicting data, effectively avoiding problems such as incorrect electricity bill collection and inaccurate repair addresses caused by address conflicts.
[0098] This invention features a hierarchical design tailored to the power industry, breaking through the limitations of a five-level system. For the first time, this invention constructs a seven-level address system for the power industry: "Province-City-District / County-Street-Neighborhood Committee-Road-Community / Village." The sixth level, "Road," is adapted for fault repair route planning (repair personnel can quickly locate fault areas using road information), while the seventh level, "Community / Village," is adapted for scenarios such as centralized meter reading (meter readers plan routes by community group) and charging pile deployment (precise site selection based on community electricity demand). Compared to existing technologies that only cover the first five levels of addresses, this invention significantly enhances the adaptability to power business scenarios, providing stronger data support for the development of power business services.
[0099] This invention solves the technical problems of limited data sources, missing hierarchical structures, and poor data quality in existing power industry address data.
[0100] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for constructing a seven-level address standard library for the power industry, characterized in that, The method includes: S1. Using the data acquisition layer, obtain the top five levels of administrative address data and the associated data of the five levels of administrative addresses, and process them to obtain multi-source address data, including: codes, government platform address data, electricity address data, and map vendor address data; S2. Using the data preprocessing layer, address segmentation, NLP, and regular expressions are employed to clean, reduce noise, and perform structure transformation operations on the multi-source address data to obtain preprocessed multi-source address data. S3. By utilizing the hierarchical construction layer and integrating government platform and map merchant data, based on the preprocessed multi-source address data, supplementing the sixth-level road and seventh-level community addresses, a seven-level address framework is constructed. S4. Using the address integration layer, establish a priority comparison model for the top five levels of addresses, obtain and integrate data from the Ministry of Civil Affairs and the National Bureau of Statistics to form basic data for the top five levels of standard addresses. S5. Utilize the address improvement layer, integrate government platform data, and perform initial construction of the first seven-level address database based on the seven-level address framework. By comparing the electricity address data with the preprocessed multi-source address data, perform model mutual verification to obtain the improved level six and level seven addresses. S6. Using the standard library output layer, based on the improved Level 6 and Level 7 addresses, output the Level 7 standard address library for the power industry.
2. The method for constructing a seven-level address standard library for the power industry according to claim 1, characterized in that, In step S2, non-hierarchical information is removed and the description is unified according to the hierarchical name extraction rules, thus completing the preprocessing operation of the electricity address data. The confidence weights of each data source at different levels and in different regions are quantified, and the confidence of each data source is quantitatively evaluated to construct a dimension scoring-weight calculation evaluation model. In the operation of quantitatively assessing the confidence of each data source, single-dimensional score statistics are performed, and scores for each dimension are calculated through field sampling. A weight calculation model is constructed, and the entropy weight method is used to determine the weight of each dimension. The comprehensive confidence weight of the data source is calculated by combining the single-dimensional scores.
3. The method for constructing a seven-level address standard library for the power industry according to claim 2, characterized in that, In step S4, the comprehensive confidence weight of the data source is calculated. W During the process, the entropy values of each dimension are calculated. This reflects the contribution of dimensional information: , Among them, the number of data sources , For the first The data source in the first The percentage of scores for each dimension; According to the dimensional entropy value Calculate dimensional weights : ; Calculate the overall confidence weight of the data source : In the formula, For the first The first data source Dimensional standardized score.
4. The method for constructing a seven-level address standard library for the power industry according to claim 1, characterized in that, In step S4, based on the priority comparison model, the basic data of the first five standard addresses are integrated; the data of the Ministry of Civil Affairs and the National Bureau of Statistics are aligned by field to obtain aligned data; and data coverage comparison and data difference comparison are performed.
5. The method for constructing a seven-level address standard library for the power industry according to claim 1, characterized in that, In step S4, when both the Ministry of Civil Affairs and the National Bureau of Statistics have address data for the same administrative unit, a difference comparison operation is performed. Based on preset conditions, the final data is determined; NLP semantic similarity calculation is used to detect name differences; and encoding difference processing is performed.
6. The method for constructing a seven-level address standard library for the power industry according to claim 1, characterized in that, In step S5, multi-source mutual verification is performed based on name comparison. Taking the initial version of the seven-level address database as a benchmark, the model mutual verification operation is completed by comparing the consistency of names at the same level and verifying the rationality of the hierarchical relationship. The consistency score is calculated by combining the confidence weight of the data source, and the problems of the initial version of the seven-level address database are identified and corrected.
7. The method for constructing a seven-level address standard library for the power industry according to claim 6, characterized in that, In the process of comparing names at the same level, name similarity is calculated by fusing edit distance similarity and Jaccard similarity, covering character differences and semantic associations; The edit distance similarity is calculated using the following logic: In the formula, This is the name of the initial version of the library to be verified. To supplement the names of data sources at the same level, To edit distance, The length of the name in characters; Calculate the Jaccard similarity using the following logic: ; Using the following logic, calculate the fusion similarity and take the average of the two as the final name similarity at the same level: 。 8. A method for constructing a seven-level address standard library for the power industry according to claim 6, characterized in that, In calculating the confidence-weighted consistency score, the overall consistency score for names at the same level is calculated by combining the confidence weights of each data source. : In the formula, For the first The similarity between each data source and the initial version of the library at the same level of integration. Assign confidence weights to the corresponding data sources; During the process of verifying the rationality of the hierarchical relationship, the hierarchical relationship is represented and matched; and the relationship consistency score is calculated. Using the following logic, combined with the confidence weight of the data source, the consistency score of the relationship is calculated. : In the formula, It is an indicator variable.
9. The method for constructing a seven-level address standard library for the power industry according to claim 1, characterized in that, In step S5, address database omissions are filled based on high-confidence names. This is achieved by comparing names from multiple sources with those from the initial seven-level address database to identify addresses not covered in the initial database, and then filling these omissions based on names from high-confidence data sources. Specifically, a reverse matching operation is used, comparing name combinations from multiple sources with the seven-level name combinations from the initial seven-level address database to identify missing addresses. For these missing addresses, the hierarchical relationships are filled in. Finally, missing address encoding and assignment operations are performed, and the filling effect is evaluated.
10. A system for constructing a seven-level address standard library for the power industry, characterized in that, The system includes: The data acquisition layer is used to acquire the top five levels of administrative address data and the associated data of the five levels of administrative addresses, and processes them to obtain multi-source address data, including: codes, government platform address data, electricity address data, and map vendor address data; The data preprocessing layer is used to perform cleaning, noise reduction, and structure transformation operations on the multi-source address data using address segmentation, NLP, and regular expressions to obtain preprocessed multi-source address data. The data preprocessing layer is connected to the data acquisition layer. The hierarchical construction layer is used to integrate government platform and map provider data. Based on the preprocessed multi-source address data, it supplements the sixth-level road and seventh-level community addresses to construct a seven-level address framework. The hierarchical construction layer is connected to the data preprocessing layer. The address integration layer is used to establish a priority comparison model for the top five levels of addresses, acquire and integrate data from the Ministry of Civil Affairs and the National Bureau of Statistics to form the basic data for the top five levels of standard addresses. The address integration layer is connected to the hierarchical construction layer. The address improvement layer is used to integrate government platform data. Based on the seven-level address framework, it performs an initial construction operation on the first seven-level address database. By comparing the electricity address data with the preprocessed multi-source address data, it performs model mutual verification operation to obtain the improved level six and level seven addresses. The address improvement layer is connected to the address integration layer. The standard library output layer is used to output a power industry level 7 standard address library based on the improved level 6 and level 7 addresses. The standard library output layer is connected to the address improvement layer.
Citation Information
Patent Citations
Address data management method based on deep learning
CN119669491A