A user address consistency research and judgment method and system
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-16
- Publication Date
- 2026-08-11
AI Technical Summary
用户地址信息存在录入不规范、变更滞后、空间指向不准确等问题;
通过融合地址文本、北斗定位、电气负荷三类异构数据,并构建时空联合演化模型,实现了多源信息的交叉验证,有效克服了单一数据源不准确或静态分析方法的局限性,显著提升了地址一致性研判的准确性和对复杂场景的适应能力;
Smart Images

Figure CN122548592A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent sensing and information processing technology of power systems, and particularly relates to a user address consistency judgment method system. Background Technology
[0002] With the continuous improvement of the digitalization and intelligence level of power distribution networks, the number of users in low-voltage distribution areas is growing rapidly, and the power consumption structure is becoming increasingly complex. Existing power distribution management systems mainly rely on user-reported addresses, manually maintained ledger information, and static topology relationships for management. However, in actual operation, the following problems are commonly encountered: The user address information has issues such as non-standard input, delayed changes, and inaccurate spatial pointing; Some users have relocated, illegally connected, or incorrectly connected their metering devices, resulting in inaccurate topology information for the transformer area that does not match the records. Existing address verification methods are mostly based on simple static address text comparison or load characteristic analysis at a single point in time, which makes it difficult to effectively capture and reflect the dynamic changes and correlations of users' electricity consumption behavior and their location in the time and space dimensions; Anomaly identification often relies on expert experience to set fixed thresholds, lacking a unified technical model that can integrate multi-source information, adaptively adjust, and be generalized.
[0003] Therefore, there is an urgent need for a comprehensive analysis method that can deeply integrate spatial location information, temporal evolution sequence and electricity consumption behavior characteristics, so as to achieve accurate and dynamic judgment of user address consistency and automatic anomaly identification, and provide a reliable data foundation for the lean management of distribution networks. Summary of the Invention
[0004] This invention aims to overcome the shortcomings of existing technologies and provide a method and system for judging user address consistency. This method constructs a three-dimensional verification system comprising "archive address semantic space - BeiDou physical space - electricity consumption behavior feature space" and introduces a time evolution dimension to achieve intelligent identification of user address anomalies and dynamic verification of transformer area topology.
[0005] In a first aspect, the present invention provides a method for determining user address consistency, comprising: Acquire user profile address information, BeiDou positioning timing information, and electrical load timing data within the target area; The file address information is parsed in two channels to generate a multi-level address semantic fingerprint vector. The address semantic distance between users is calculated based on the address semantic fingerprint vector. Combined with the preset BeiDou spatial aggregation constraint and the preset electrical aggregation constraint, the first user grouping result and the first group confidence degree based on address semantics are obtained. Based on the electrical load time series data, the multi-domain electrical distance between users in multiple time windows is calculated. The multi-domain electrical distance, address semantic distance, BeiDou spatial distance and preset structural prior soft constraints are combined and fused through reliability-driven adaptive weights to generate the final distance matrix under spatiotemporal joint evolution. Clustering is performed based on the final distance matrix to obtain the second user grouping results and the second grouping confidence based on load characteristics; When the first user grouping result based on address semantics, the second user grouping result based on load characteristics, and the spatial aggregation result based on BeiDou positioning information conflict or show deviations that do not conform to the preset evolution law within a continuous time window, it is determined that the user address consistency is abnormal, and the abnormal evidence chain and abnormal confidence score are output.
[0006] Secondly, the present invention provides a user address consistency assessment system, comprising: The acquisition module is configured to acquire user profile address information, BeiDou positioning time sequence information, and electrical load time sequence data within the target area; The parsing module is configured to perform dual-channel parsing on the file address information, generate multi-level address semantic fingerprint vectors, calculate the address semantic distance between users based on the address semantic fingerprint vectors, and combine preset BeiDou spatial aggregation constraints and preset electrical aggregation constraints to obtain the first user grouping result and the first grouping confidence based on address semantics. The fusion module is configured to calculate the multi-domain electrical distance between users within multiple time windows based on the electrical load time series data, and combine the multi-domain electrical distance, address semantic distance, BeiDou spatial distance and preset structural prior soft constraints, and fuse them through reliability-driven adaptive weights to generate the final distance matrix under spatiotemporal joint evolution. The clustering module is configured to perform clustering based on the final distance matrix to obtain a second user grouping result and a second grouping confidence level based on load characteristics. The determination module is configured to determine that the user address consistency is abnormal when the first user grouping result based on address semantics, the second user grouping result based on load characteristics, and the spatial aggregation result based on BeiDou positioning information conflict or show deviations that do not conform to the preset evolution law within a continuous time window, and output the abnormal evidence chain and abnormal confidence score.
[0007] Thirdly, an electronic device is provided, comprising: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the steps of the user address consistency assessment method according to any embodiment of the present invention.
[0008] Fourthly, the present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein when the program instructions are executed by a processor, the processor performs the steps of the user address consistency assessment method of any embodiment of the present invention.
[0009] The user address consistency assessment method and system proposed in this application have the following specific benefits: By integrating three types of heterogeneous data—address text, BeiDou positioning, and electrical load—and constructing a spatiotemporal joint evolution model, cross-validation of multi-source information was achieved. This effectively overcomes the limitations of inaccurate single data sources or static analysis methods, significantly improving the accuracy of address consistency assessment and adaptability to complex scenarios. By using the time-series characteristics of electrical loads for evolution analysis, it is possible to automatically identify changes in topology caused by user migration, unauthorized connections, incorrect connections, etc., thereby realizing dynamic and automatic verification of transformer area topology and reducing reliance on manual maintenance. The system introduces a reliability-driven adaptive weight fusion mechanism, conflict-driven alias learning, and intelligent judgment rules based on consistency violation, enabling the entire judgment process to have self-learning and adaptive capabilities, reducing reliance on preset thresholds and expert experience, and achieving automatic and accurate identification and location of anomalies. Attached Figure Description
[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 A flowchart illustrating a user address consistency assessment method according to an embodiment of the present invention; Figure 2 This is a structural block diagram of a user address consistency assessment system provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation
[0012] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0013] Please see Figure 1 The diagram shows a flowchart of a user address consistency assessment method according to this application.
[0014] like Figure 1 As shown, the user address consistency assessment method specifically includes the following steps: Step S101: Obtain the user's file address information, BeiDou positioning time sequence information, and electrical load time sequence data within the target area.
[0015] Step S102: Perform dual-channel parsing on the file address information to generate multi-level address semantic fingerprint vectors. Calculate the address semantic distance between users based on the address semantic fingerprint vectors. Combine the preset BeiDou spatial aggregation constraints and preset electrical aggregation constraints to obtain the first user grouping result and the first grouping confidence level based on address semantics.
[0016] In this step, the address string in the archive address information is segmented at the character level through the character channel, and the unification of Arabic numerals and Chinese numerals, the unification of full-width and half-width characters, and the normalization of preset common abbreviations are completed. Through the place name channel, based on the preset administrative division tree and place name database, the address string is parsed into a candidate segment sequence that constitutes a hierarchical structure. The hierarchical structure is the structure of province, city, county, township, village, natural group, building, unit, and room number. The optimal parsing sequence is selected from the candidate segment sequence through the minimum conflict path search algorithm. For each level field obtained after parsing, a composite fingerprint of the corresponding level field is generated. The composite fingerprint includes at least a hash fingerprint of the field text, a phonetic near-sound code fingerprint for handling homophones or misspellings, and a sequence-sensitive n-gram signature for character order characteristics within the field. The composite fingerprints of all level fields are combined in hierarchical order to form the multi-level address semantic fingerprint vector.
[0017] Further, based on the address semantic fingerprint vectors, calculate the fingerprint similarity between any two users at each level; adaptively adjust the weight of the level similarity in the overall similarity calculation according to the structural integrity of each level field; comprehensively calculate the weighted similarities of each level to obtain the address semantic distance; construct a user similarity graph with the address semantic distance, and use the Beidou spatial aggregation constraint and the electrical aggregation constraint as the constraint edges that must be connected or prohibited from being connected in the user similarity graph; perform connected component analysis or constraint clustering algorithm on the user similarity graph with electrical aggregation constraints to obtain the first user grouping result; calculate the confidence of the first grouping based on the dispersion degree or constraint satisfaction degree of the address semantic distance within the group.
[0018] In a specific embodiment, the address parsing adopts the parallel cooperation of the "character channel" and the "place name channel". The character channel first performs character-level preprocessing on the original address string, including: unifying Chinese numbers (such as "two zero one") and Arabic numbers; converting full-width characters (such as ",", "") to half-width; and normalizing common abbreviations (such as simplifying "Limited Company" to "Company", "Building No. X" to "Building X") to eliminate expression differences. The place name channel is based on a pre-constructed knowledge base that includes the national administrative division tree (province, city, county, township, village) and a large number of standard place names (road, community, building name). This channel matches the preprocessed address string with the knowledge base to generate all possible candidate fragment sequences that conform to the hierarchical logic of "province-city-county-township-village-natural group-building-unit-room number". For example, "Room 303, Building 5, Yard No. 27, Zhongguancun Street, Haidian District, Beijing" may be parsed into multiple candidate paths such as {province: Beijing, city: Beijing, district: Haidian District, street: Zhongguancun Street, road: Zhongguancun Street, house number: No. 27, building: Building 5, room number: 303}.
[0019] To select the most correct one from multiple candidate paths, the system uses the minimum conflict path search algorithm. This algorithm regards each candidate parsing path as a graph node, and the differences between paths (such as field overlap, logical contradiction) constitute conflict edges. The goal of the algorithm search is to find a path that covers all valid address components and has the minimum total internal conflict as the optimal parsing sequence, so as to effectively handle complex addresses such as "XX Road YY No. ZZ Community" that may have nesting or ambiguity.
[0020] After parsing, a composite fingerprint is generated for each level field in the optimal sequence (such as "Haidian District" and "Building No. 5"). The fingerprint consists of three parts: (1) Hash fingerprint: The standard text of the field is calculated using the MD5 or SHA-256 algorithm for fast and accurate matching; (2) Phonetic fingerprint: The field text is converted into pinyin (such as "Haidian District" is converted into "hai dian qu"), and then its first letter ("hdq") can be taken to tolerate homophones, pinyin input errors or typos caused by dialects; (3) Sequence-sensitive n-gram signature: The field text is slid segmented into 2-gram or 3-gram (such as the 2-gram of "Building No. 5" is {"No. 5", "No. Building"}), and a signature is generated to capture the character order features inside the field and prevent the disordered order of "Building No. 5" from being misjudged as similar. Finally, the composite fingerprints of all hierarchical fields are concatenated in a fixed hierarchical order of "province, city, county..." to form a high-dimensional, multimodal address semantic fingerprint vector, which serves as the unique digital identity of the user's address.
[0021] After obtaining the address semantic fingerprint vectors of all users, the system calculates the address semantic distance between any two users. The calculation is not a simple vector cosine similarity, but rather employs a hierarchical weighted strategy: first, the similarity of the fingerprints of the two users at each same level is calculated (combining hash, phonetic similarity codes, and n-grams); then, weights are adaptively assigned based on the structural completeness of that level (e.g., a higher weight is given if the "building" field information is complete, and a lower weight or even zero weight is given if it is missing); finally, the weighted similarities of all levels are aggregated to obtain the comprehensive distance. This method makes the distance metric more robust to partial missing address information or errors at minor levels.
[0022] To improve grouping accuracy, the system introduces two types of prior constraints to construct a user similarity graph (with users as nodes and address semantic distance as edge weights): BeiDou spatial clustering constraints: If the BeiDou positioning coordinates of two users are very close in physical space (e.g., the horizontal distance is less than 10 meters), a strong constraint edge of "must connect" is added to the graph, forcing them to be prioritized for merging in future groupings; conversely, if the distance is extremely far (e.g., more than 1 kilometer) and there is no other evidence of correlation, a "do not connect" constraint may be added.
[0023] Electrical aggregation constraint: Import known topological relationships from the power distribution management system, such as user pairs under the same meter box, the same incoming line, or the same transformer, and add soft constraint edges that "should belong to the same group", assigning a higher initial connection weight.
[0024] Building upon this, the system performs connected component analysis or constrained spectral clustering algorithms on the constrained similarity graph. The algorithms aim to minimize the semantic distance between addresses within a group while satisfying these spatial and electrical constraints as much as possible. The final output is the first user grouping result based on address semantics, grouping users with similar addresses and spatial or electrical connections into the same group. Simultaneously, the system evaluates the first grouping confidence of each group, calculated based on: the average cohesion of the address semantic vectors of members within the group (lower dispersion indicates higher confidence), and the proportion of the grouping result satisfying constraints such as "must be connected" and "should belong to the same group" (higher satisfaction indicates higher confidence).
[0025] In summary, by employing a dual-channel collaborative parsing approach involving characters and place names, combined with minimum conflict path search, the accuracy and robustness of structured parsing for complex, non-standard address text are significantly improved. This effectively addresses real-world issues such as address abbreviations, aliases, out-of-order entries, and typos, providing a high-quality semantic foundation for subsequent analysis. Secondly, a multimodal composite fingerprint vector (integrating hash, phonetic similarity, and sequence features) is innovatively proposed, achieving refined and noise-resistant representation of address semantics from text to high-dimensional vectors, making address similarity calculations more comprehensive and accurate. Thirdly, the BeiDou spatial clustering and electrical topology soft constraints are creatively introduced during the clustering process, upgrading pure text analysis to multi-source information fusion decision-making. This ensures that grouping results are not only based on "textual similarity" but also consider "physical proximity" and "electrical correlation," significantly improving grouping accuracy even with incomplete or slightly inaccurate user address registrations. This lays a reliable foundation for subsequent cross-spatial consistency comparisons with load feature groupings. Fourth, the output group confidence level provides a quantifiable reliability indicator for the results, helping maintenance personnel to focus on low-confidence and high-uncertainty groups for key verification, thereby improving the efficiency and pertinence of the overall analysis work.
[0026] Step S103: Based on the electrical load time series data, calculate the multi-domain electrical distance between users in multiple time windows. Combine the multi-domain electrical distance, address semantic distance, BeiDou spatial distance and preset structural prior soft constraints, and fuse them through reliability-driven adaptive weights to generate the final distance matrix under spatiotemporal joint evolution.
[0027] In this step, the electrical load time-series data is captured using a sliding time window. Within each time window, the multi-dimensional electrical characteristic distance between any two users is calculated. The multi-dimensional distance includes at least: the distance based on the correlation coefficient and dynamic time warping of the voltage sequence, the distance based on the correlation coefficient of the active power and reactive power sequences, and the vector distance based on the current harmonic spectrum. For three-phase users, an optimized assignment algorithm is used to find the minimum cost match of the multi-dimensional electrical characteristic distance in phases A, B, and C, which is taken as the multi-domain electrical distance between users within the time window.
[0028] Furthermore, data source reliability indices are constructed for archive address information, BeiDou positioning information, and electrical load time-series data, respectively. For any two users, within each time window, based on the data source reliability indices and the degree of cross-domain consistency conflict between multi-domain electrical distance, address semantic distance, and BeiDou spatial distance, the first fusion weight of multi-domain electrical distance, the second fusion weight of address semantic distance, and the third fusion weight of BeiDou spatial distance are dynamically calculated. Based on the first fusion weight, the second fusion weight, and the third fusion weight, the multi-domain electrical distance, address semantic distance, and BeiDou spatial distance are weighted and summed to obtain the preliminary fusion distance under the time window. The preliminary fusion distance of each time window is standardized by quantiles, and the structural prior soft constraint representing the prior knowledge of the distribution network topology is introduced as a penalty term to obtain the fusion distance matrix. The standardized and penalized fusion distance matrix is globally corrected based on the shortest path closure algorithm to make the fusion distance matrix satisfy the triangle inequality of the metric space, thereby generating the final distance matrix under the spatiotemporal joint evolution.
[0029] In one specific embodiment, a sliding time window mechanism is first used to segment and analyze continuous electrical load time-series data. For example, the window length is set to 24 hours and the sliding step size is 1 hour to capture the continuous evolution of daily load patterns. Within each time window t, for any two users i and j, their multi-dimensional electrical characteristic distance is calculated. This distance integrates similarity measures from multiple electrical physical domains. Voltage Domain Distance: Calculates the Pearson correlation coefficient and dynamic time warping distance (DTW) of two user voltage sequences. The correlation coefficient measures the similarity of curve shapes, while the DTW distance effectively aligns and compares curves with slight offsets or stretching on the time axis. Combining the two provides a comprehensive assessment of the synchronicity and shape differences of voltage changes.
[0030] Power domain distance: Calculate the correlation coefficient distance (e.g., 1 - correlation coefficient) between the active power series and reactive power series of two users. This reflects the coordinated pattern of users' electricity consumption and power factor changes, and is key to determining whether their electricity consumption behavior originates from the same source.
[0031] Harmonic domain distance: Extract the amplitude of characteristic subharmonics (such as the 3rd, 5th, and 7th harmonics) from the current waveforms of two users, construct the harmonic spectrum vector, and calculate their Euclidean or cosine distance. Harmonic characteristics are strongly correlated with the nonlinear electrical equipment (such as power electronic devices) inside the user's device, and can serve as a "fingerprint" to distinguish users with different electricity consumption types.
[0032] For three-phase users, their load curves need to be compared with the three phases (A, B, and C) of the main meter in the distribution area. The system uses an optimized assignment algorithm (such as the Hungarian algorithm) to best match the user's three-phase load curve with the three phases of the main meter, in order to find the phase assignment scheme with the minimum total cost (i.e., the sum of the three-phase electrical characteristic distances). This minimum total cost is the core component of the final multi-domain electrical distance between the user and the main meter within this time window. The electrical distance between users is calculated by comparing their respective (phase-assigned) load characteristics.
[0033] For each type of data, data source reliability metrics are constructed: structural integrity of address information (scored based on field missing rate and hierarchical logical consistency), reliability of BeiDou positioning (combining positioning accuracy, timestamp freshness, and historical coordinate stability), and stability of electrical load (calculated based on data missing rate, outlier removal ratio, and fluctuation coefficient). These metrics quantify the quality of each data source within the current analytical context.
[0034] For any user pair (i,j) within the time window t, the system has three basic distances: (Multi-domain electrical distance), (Address semantic distance, from step S102) (BeiDou space Euclidean distance). The core of the fusion is the dynamic calculation of three adaptive weights. , , .
[0035] Weight calculation logic: The weight of each distance is positively correlated with the reliability of its data source (the higher the reliability, the larger the base weight value), and negatively correlated with the degree of cross-domain consistency conflict. The degree of conflict is calculated by comparing the relative magnitude and trend between each pair of the three distances (e.g., if...). The small size indicates a high degree of similarity in electricity consumption behavior, but A large value indicates a large physical distance, which suggests a high degree of conflict between electrical distance and spatial distance, and the weights of both may be appropriately reduced. The dynamic weighting model ensures that evidence in this dimension is strengthened when the data quality is high and they corroborate each other, while reducing its impact when the data quality is poor or there are contradictions, thus achieving robust fusion.
[0036] The initial fusion distance under time window t is obtained by weighted summation based on dynamic weights. .
[0037] To eliminate differences in dimensions and benchmarks between different time windows, the preliminary fusion distance matrices for all time windows are standardized using quantiles. Subsequently, a structural prior soft constraint matrix P is introduced as a penalty term. Matrix P is constructed based on distribution network topology knowledge: for example, if users i and j are known to be located on the same feeder or under the same transformer, P(i,j) is set to a negative value (reducing the penalty and encouraging proximity); if there is no explicit connection and the distance is far, it is set to a positive value or zero. The standardized distance matrix is then superimposed with the penalty matrix.
[0038] Finally, to ensure the generated distance matrix possesses sound mathematical properties to support high-quality clustering, the system employs a shortest path closure algorithm (such as the Floyd-Warshall algorithm) to globally correct the superimposed matrix. This algorithm calculates the shortest path distance between all pairs of nodes in the graph using the current distance as the edge weight, and replaces the original direct distances with this shortest path distance. This operation forces the distance matrix to satisfy the triangle inequality (i.e., for any three points, the sum of any two sides is greater than the third side), thus forming a metric space. The final distance matrix D_final(t) generated after this correction not only integrates multi-source spatiotemporal information, but its own metric characteristics also make the "distance" relationship between users more geometrically intuitive, providing ideal input for subsequent density-based clustering analysis.
[0039] Step S104: Clustering is performed based on the final distance matrix to obtain the second user grouping results and the second grouping confidence based on load characteristics.
[0040] In this step, within each independent time window, the final distance matrix corresponding to the time window is used as input, and a density-based clustering algorithm is employed to cluster users, resulting in instantaneous user groups within that time window. The frequency at which any two users are assigned to the same instantaneous group within all time windows is counted to construct a co-cluster matrix. Consistent clustering is then performed on the co-cluster matrix, aggregating users with co-cluster frequencies higher than a stability threshold into the same stable group, thus obtaining the second user grouping result. Based on the intra-cluster consistency and cross-time window stability of the co-cluster matrix, the confidence level of the second group is calculated.
[0041] In one specific embodiment, for each independent time window t, the system uses its corresponding final distance matrix D_final(t) as input and employs a density-based clustering algorithm (DBSCAN) for clustering. The DBSCAN algorithm does not require pre-specifying the number of clusters, can automatically discover clusters of arbitrary shapes, and can effectively identify noise points (i.e., outliers that do not belong to any stable cluster). This is very suitable for handling real-world scenarios with uneven distribution of power distribution users and irregular cluster shapes. The system adaptively sets the neighborhood radius ε and the minimum number of points MinPts parameters based on data characteristics (such as the median of the distance distribution), enabling the algorithm to discover tightly clustered user groups in the distance space of that time window. After clustering, each user is labeled as belonging to a certain "instantaneous cluster" or labeled as "noise." This process generates a set of instantaneous user groups C(t) = {C1(t), C2(t), ...} for each time window t, reflecting the local clustering state based on comprehensive spatiotemporal behavioral characteristics at that moment.
[0042] To extract stable grouping patterns from a series of potentially fluctuating instantaneous groupings, the system counts the frequency with which any two users are assigned to the same instantaneous cluster within all time windows (e.g., all windows over the past 30 days). Specifically, an N×N matrix M (where N is the total number of users) is initialized. For each pair of users (i, j), all time windows are iterated. If i and j are assigned to the same cluster by DBSCAN within the same window (and neither is a noise point), the count value of M(i, j) is incremented by 1. Finally, the count value is divided by the total number of time windows to obtain the co-cluster frequency matrix (or co-occurrence matrix). The element freq(i, j) of this matrix represents the frequency with which users i and j exhibit highly similar comprehensive behaviors (address, space, electrical) across all observation periods. The higher the frequency value, the more stable the relationship between the two and the more likely they belong to the same inherent group.
[0043] After obtaining the co-cluster frequency matrix, the system treats it as a new similarity metric (higher frequency indicates greater similarity) and performs consistent clustering on it. An agglomerative hierarchical clustering algorithm can be used here. The algorithm starts by treating each user as a separate cluster, then iteratively merges the two most similar clusters (i.e., those with the highest co-cluster frequencies) until the similarity between all clusters falls below a preset stability threshold (e.g., 0.7, indicating that users must behave consistently over more than 70% of the time windows). This threshold determines the "strictness" of the grouping. The resulting cluster partition represents a set of users with highly consistent behavioral patterns across the time dimension, serving as the second user grouping result G_load={G1, G2, ...} based on the spatiotemporal evolution characteristics of load.
[0044] Intra-cluster consistency: Calculate the average co-cluster frequency of all user pairs within group Gk. A higher average frequency indicates better synchronization of behavior among group members across all time windows, resulting in higher consistency.
[0045] Stability across time windows: This analyzes the regularity of users constituting Gk being grouped into the same instantaneous cluster across historical time windows. For example, it calculates the variance of the simultaneous occurrence of this group of users across different time windows. The smaller the variance, the more stable the grouping pattern is, and the less affected it is by short-term fluctuations.
[0046] Through a weighted model (e.g., C_load = α) (Intra-cluster average frequency) + β (1 - stability variance), where α and β are harmonic weights. These two indicators are combined to output a confidence score between 0 and 1. High confidence groups are considered highly reliable and have clear patterns, while low confidence groups suggest that they may contain users with variable behavioral patterns or who are on the margins and require further attention.
[0047] Step S105: When the first user grouping result based on address semantics, the second user grouping result based on load characteristics, and the spatial aggregation result based on BeiDou positioning information conflict or show deviations that do not conform to the preset evolution law within a continuous time window, it is determined that the user address consistency is abnormal, and the abnormal evidence chain and abnormal confidence score are output.
[0048] In this step, the conflict is determined as follows: when the same user is assigned to different groups in the first user grouping result and the second user grouping result, and the distance between its BeiDou positioning coordinates and the spatial aggregation center of any of its groups exceeds a preset threshold, it is determined that there is a semantic-electrical-spatial consistency conflict. Dynamic anomaly determination is based on tracking the evolution of a user's group within a continuous time window. When a user's trend of leaving the original group or drifting to other groups does not conform to the preset normal evolution pattern learned from historical data, it is determined to be a dynamic evolution anomaly.
[0049] In one specific embodiment, the chain of evidence is a data structure containing multiple fields, the core fields of which are as follows: Anomaly types: Static conflict (inconsistency in three spaces) / Dynamic evolution anomaly (grouping jump / drift).
[0050] Conflict details: For static conflicts, record. , , , Distance to each center, spatial consistency threshold .
[0051] Dynamic details: For dynamic anomalies, record the anomaly start time window and the original group. Target group G_new (if identifiable), evolution trend chart or quantitative indicator (such as drift speed).
[0052] Data source quality: structural integrity associated with the user profile address, reliability of BeiDou positioning, and stability of electrical data.
[0053] Timestamp: The moment when the anomaly was detected.
[0054] Confidence score calculation: The confidence score, Score_anomaly(u), integrates the reliability of multiple pieces of evidence and is calculated using a pre-defined scoring model that considers the following factors: Group confidence weight: The group confidence (C_addr, C_load) of the group involved in the anomaly (such as G_addr(u) and G_load(u)). The higher the group confidence, the more reliable the anomaly evidence provided by that group, and the greater its positive weight contribution.
[0055] Significance of conflict / deviation: For static conflicts, calculate (actual spatial offset distance - / As a measure of spatial deviation significance; for dynamic anomalies, the slope of the evolution trend or the magnitude of the difference in features before and after the jump is calculated. The higher the significance, the higher the confidence level.
[0056] Duration over time: The number of time windows during which the anomalous pattern persists. The longer the duration, the lower the randomness and the higher the confidence level.
[0057] Data source reliability: The weighted average of user data source reliability metrics (structural integrity, location reliability, electrical stability). The higher the overall quality of the data source, the more reliable the judgment result.
[0058] The final score is usually normalized to the [0, 1] interval, for example, using a weighted summation formula: Score_anomaly(u) = w1 F(group confidence level) + w2 F(significance) + w3 F (persistence) +w4 F (Data source reliability) Where F() is the normalization function, and w1, w2, w3, and w4 are adjustable weights. A high score (e.g., >0.8) indicates conclusive evidence of anomalies and suggests prioritizing its handling; a low score (e.g., 0.4-0.7) indicates the presence of anomalies but requires further investigation.
[0059] The final anomaly assessment results (list of anomaly users, evidence chain, confidence score) are pushed to the distribution network management system's visualization platform or work order system. Maintenance personnel can prioritize high-risk anomalies based on confidence scores and conduct precise on-site verification using detailed information in the evidence chain (such as spatial location and electricity consumption curve comparison). The verification results can be fed back to the system for model optimization (such as updating normal evolution patterns and adjusting thresholds), forming a closed-loop business process of "monitoring-assessment-handling-optimization."
[0060] In summary, the method of this application, at the level of data fusion and representation, overcomes the structural challenges of non-standard and ambiguous address text through innovative dual-channel address parsing and multimodal semantic fingerprint vector construction, achieving high-precision digital expression of address semantics. At the same time, it introduces a reliability-driven adaptive weight fusion mechanism, and for the first time incorporates the quality assessment and cross-domain consistency dynamics of three types of heterogeneous data—archival addresses, BeiDou positioning, and electrical loads—into the decision-making process, constructing a spatiotemporal joint evolution model that can self-evaluate and resist interference, fundamentally improving the intelligence and robustness of multi-source information fusion. Secondly, in terms of analytical and identification capabilities, this method achieves a leap from static comparison to dynamic perception: On the one hand, through the three-space consistency conflict criterion of "address semantics-physical space-electrical characteristics," it achieves high-precision and automated identification of static anomalies such as file errors and meter misconnections, significantly reducing misjudgments caused by inaccurate single data sources; on the other hand, it pioneers time-series evolution analysis based on user behavior trajectories, extracting steady-state groups through two-stage clustering and monitoring dynamic drift, enabling it to keenly capture gradual address changes or hidden illegal electricity use behaviors that traditional methods cannot detect, achieving an upgrade in proactive defense capabilities from post-event verification to in-event early warning. Finally, in terms of output and practical application, this method not only outputs anomaly conclusions but also simultaneously generates structured anomaly evidence chains and quantified confidence scores. The evidence chains clearly reveal the complete logic of "who, when, and why the conflict occurs," providing precise navigation for on-site verification and greatly improving operational efficiency; while the confidence scores provide intelligent priority ranking for massive alarms, achieving optimized allocation of operational resources. In summary, this invention comprehensively improves the accuracy, depth, intelligence level, and decision support value of user address consistency assessment, providing strong core technical support for dynamic topology verification, lean line loss management, and anti-theft and anti-violation investigation of distribution networks.
[0061] Please see Figure 2 The diagram shows a structural block diagram of a user address consistency assessment system according to this application.
[0062] like Figure 2 As shown, the user address consistency assessment system 200 includes... The acquisition module 210 is configured to acquire user file address information, BeiDou positioning time series information, and electrical load time series data within the target area; the parsing module 220 is configured to perform dual-channel parsing on the file address information to generate multi-level address semantic fingerprint vectors, calculate the address semantic distance between users based on the address semantic fingerprint vectors, and combine preset BeiDou spatial aggregation constraints and preset electrical aggregation constraints to obtain the first user grouping result and the first grouping confidence level based on address semantics; the fusion module 230 is configured to calculate the multi-domain electrical distance between users within multiple time windows based on the electrical load time series data, and combine the multi-domain electrical distance and address semantic distance... The distance between the BeiDou and other satellites, along with the preset structural prior soft constraints, are fused using reliability-driven adaptive weights to generate a final distance matrix under spatiotemporal joint evolution. Clustering module 240 is configured to perform clustering based on the final distance matrix to obtain a second user grouping result and a second grouping confidence score based on load characteristics. Judgment module 250 is configured to determine user address consistency anomalies when the first user grouping result based on address semantics, the second user grouping result based on load characteristics, and the spatial aggregation result based on BeiDou positioning information conflict or exhibit deviations that do not conform to preset evolutionary rules within a continuous time window, and outputs an abnormal evidence chain and an abnormal confidence score.
[0063] It should be understood that Figure 2 The modules and references described in the document Figure 1 The steps described in the text correspond to those in the method described above. Therefore, the operations, features, and corresponding technical effects described above also apply to the method described in the text. Figure 2 The various modules in the document will not be described in detail here.
[0064] In other embodiments, the present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein when the program instructions are executed by a processor, the processor performs the user address consistency assessment method in any of the above method embodiments. In one embodiment, the computer-readable storage medium of the present invention stores computer-executable instructions, which are configured as follows: Acquire user profile address information, BeiDou positioning timing information, and electrical load timing data within the target area; The file address information is parsed in two channels to generate a multi-level address semantic fingerprint vector. The address semantic distance between users is calculated based on the address semantic fingerprint vector. Combined with the preset BeiDou spatial aggregation constraint and the preset electrical aggregation constraint, the first user grouping result and the first group confidence degree based on address semantics are obtained. Based on the electrical load time series data, the multi-domain electrical distance between users in multiple time windows is calculated. The multi-domain electrical distance, address semantic distance, BeiDou spatial distance and preset structural prior soft constraints are combined and fused through reliability-driven adaptive weights to generate the final distance matrix under spatiotemporal joint evolution. Clustering is performed based on the final distance matrix to obtain the second user grouping results and the second grouping confidence based on load characteristics; When the first user grouping result based on address semantics, the second user grouping result based on load characteristics, and the spatial aggregation result based on BeiDou positioning information conflict or show deviations that do not conform to the preset evolution law within a continuous time window, it is determined that the user address consistency is abnormal, and the abnormal evidence chain and abnormal confidence score are output.
[0065] Computer-readable storage media may include a stored program area and a stored data area, wherein the stored program area may store an operating system and an application program required for at least one function; the stored data area may store data created based on the use of the user address consistency assessment system, etc. Furthermore, the computer-readable storage medium may include high-speed random access memory, and may also include memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device. In some embodiments, the computer-readable storage medium may optionally include memory remotely located relative to the processor, and these remote memories may be connected to the user address consistency assessment system via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0066] Figure 3 This is a schematic diagram of the structure of the electronic device provided in the embodiment of the present invention, such as... Figure 3 As shown, the device includes a processor 310 and a memory 320. The electronic device may also include an input device 330 and an output device 340. The processor 310, memory 320, input device 330, and output device 340 can be connected via a bus or other means. Figure 3 Taking a bus connection as an example, the memory 320 is the computer-readable storage medium described above. The processor 310 executes various server functions and data processing by running non-volatile software programs, instructions, and modules stored in the memory 320, thereby implementing the user address consistency assessment method described in the above embodiment. The input device 330 can receive input numeric or character information and generate key signal inputs related to user settings and function control of the user address consistency assessment system. The output device 340 may include a display screen or other display device.
[0067] The aforementioned electronic device can execute the method provided in the embodiments of the present invention, and has the corresponding functional modules and beneficial effects for executing the method. Technical details not described in detail in this embodiment can be found in the method provided in the embodiments of the present invention.
[0068] In one implementation, the above-described electronic device is used in a user address consistency assessment system for a client, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to: Acquire user profile address information, BeiDou positioning timing information, and electrical load timing data within the target area; The file address information is parsed in two channels to generate a multi-level address semantic fingerprint vector. The address semantic distance between users is calculated based on the address semantic fingerprint vector. Combined with the preset BeiDou spatial aggregation constraint and the preset electrical aggregation constraint, the first user grouping result and the first group confidence degree based on address semantics are obtained. Based on the electrical load time series data, the multi-domain electrical distance between users in multiple time windows is calculated. The multi-domain electrical distance, address semantic distance, BeiDou spatial distance and preset structural prior soft constraints are combined and fused through reliability-driven adaptive weights to generate the final distance matrix under spatiotemporal joint evolution. Clustering is performed based on the final distance matrix to obtain the second user grouping results and the second grouping confidence based on load characteristics; When the first user grouping result based on address semantics, the second user grouping result based on load characteristics, and the spatial aggregation result based on BeiDou positioning information conflict or show deviations that do not conform to the preset evolution law within a continuous time window, it is determined that the user address consistency is abnormal, and the abnormal evidence chain and abnormal confidence score are output.
[0069] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0070] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for judging user address consistency, characterized in that, include: Acquire user profile address information, BeiDou positioning timing information, and electrical load timing data within the target area; The file address information is parsed in two channels to generate a multi-level address semantic fingerprint vector. The address semantic distance between users is calculated based on the address semantic fingerprint vector. Combined with the preset BeiDou spatial aggregation constraint and the preset electrical aggregation constraint, the first user grouping result and the first group confidence degree based on address semantics are obtained. Based on the electrical load time series data, the multi-domain electrical distance between users in multiple time windows is calculated. The multi-domain electrical distance, address semantic distance, BeiDou spatial distance and preset structural prior soft constraints are combined and fused through reliability-driven adaptive weights to generate the final distance matrix under spatiotemporal joint evolution. Clustering is performed based on the final distance matrix to obtain the second user grouping results and the second grouping confidence based on load characteristics; When the first user grouping result based on address semantics, the second user grouping result based on load characteristics, and the spatial aggregation result based on BeiDou positioning information conflict or show deviations that do not conform to the preset evolution law within a continuous time window, it is determined that the user address consistency is abnormal, and the abnormal evidence chain and abnormal confidence score are output.
2. The user address consistency assessment method according to claim 1, characterized in that, The step of performing dual-channel parsing on the file address information to generate a multi-level address semantic fingerprint vector includes: The address string in the file address information is segmented at the character level through the character channel, and the unification of Arabic numerals and Chinese numerals, the unification of full-width and half-width characters, and the normalization of preset common abbreviations are completed. Through the place name channel, based on the preset administrative division tree and place name database, the address string is parsed into a sequence of candidate segments that constitute a hierarchical structure, which is a structure of province, city, county, township, village, natural group, building, unit, and room number; The optimal parsing sequence is selected from the candidate segment sequence using the minimum conflict path search algorithm; For each level of field obtained after parsing, a composite fingerprint of the corresponding level of field is generated. The composite fingerprint includes at least a hash fingerprint of the field text, a phonetic near-sound fingerprint for handling homophones or misspellings, and a sequence-sensitive n-gram signature for character order features within the field. The composite fingerprints of all hierarchical fields are combined in hierarchical order to form the multi-level address semantic fingerprint vector.
3. The user address consistency assessment method according to claim 1, characterized in that, The step of calculating the address semantic distance between users based on the address semantic fingerprint vector, and combining it with preset BeiDou spatial aggregation constraints and preset electrical aggregation constraints, to obtain the first user grouping result and the first grouping confidence level based on address semantics includes: Based on the address semantic fingerprint vector, calculate the fingerprint similarity of any two users at each level; The weight of hierarchical similarity in the overall similarity calculation is adaptively adjusted based on the structural completeness of each level field. The address semantic distance is calculated by combining the weighted similarity of each level; A user similarity graph is constructed based on the address semantic distance, and the BeiDou spatial clustering constraint and electrical clustering constraint are used as constraint edges in the user similarity graph that must be connected or must not be connected. Perform connected component analysis or constrained clustering algorithm on the user similarity graph with electrical aggregation constraints to obtain the first user grouping results; The confidence level of the first group is calculated based on the degree of dispersion or constraint satisfaction of the semantic distance of user addresses within the group.
4. The user address consistency assessment method according to claim 1, characterized in that, The calculation of multi-domain electrical distance between users within multiple time windows based on the electrical load time-series data includes: The electrical load timing data is captured using a sliding time window. Within each time window, calculate the multi-dimensional electrical characteristic distance between any two users. The multi-dimensional distance includes at least: the distance based on the correlation coefficient and dynamic time warping of the voltage sequence, the distance based on the correlation coefficient of the active power and reactive power sequences, and the vector distance based on the current harmonic spectrum. For three-phase users, an optimized assignment algorithm is used to find the minimum cost match of the multi-dimensional electrical feature distance on phases A, B, and C, which is used as the multi-domain electrical distance between users within the time window.
5. The user address consistency assessment method according to claim 1, characterized in that, The process of combining the multi-domain electrical distance, address semantic distance, BeiDou spatial distance, and preset structural prior soft constraints, and fusing them through reliability-driven adaptive weights to generate the final distance matrix under spatiotemporal joint evolution includes: Data source reliability indices were constructed for archive address information, BeiDou positioning information, and electrical load time series data, respectively; For any two users, within each time window, based on the data source reliability index and the degree of cross-domain consistency conflict between multi-domain electrical distance, address semantic distance, and BeiDou spatial distance, the first fusion weight of multi-domain electrical distance, the second fusion weight of address semantic distance, and the third fusion weight of BeiDou spatial distance are dynamically calculated. Based on the first fusion weight, the second fusion weight, and the third fusion weight, the multi-domain electrical distance, the address semantic distance, and the BeiDou spatial distance are weighted and summed to obtain the preliminary fusion distance under the time window. The initial fusion distance for each time window is standardized by quantiles, and the structural prior soft constraint representing the prior knowledge of the distribution network topology is introduced as a penalty term to obtain the fusion distance matrix. The standardized and penalized fusion distance matrix is globally corrected based on the shortest path closure algorithm to ensure that the fusion distance matrix satisfies the triangle inequality of the metric space, thereby generating the final distance matrix under the spatiotemporal joint evolution.
6. The user address consistency assessment method according to claim 1, characterized in that, The clustering based on the final distance matrix to obtain the second user grouping results and the second grouping confidence based on load characteristics includes: In each independent time window, the final distance matrix corresponding to the time window is used as input, and a density-based clustering algorithm is used to perform clustering to obtain the instantaneous user groups under the time window; Count the frequency at which any two users are assigned to the same instantaneous group within all time windows, and construct a co-cluster matrix; Perform consistent clustering on the co-cluster matrix to aggregate users whose co-cluster frequency is higher than the stability threshold into the same stable group, and obtain the second user grouping result; The confidence level of the second group is calculated based on the intra-cluster consistency and cross-time window stability of the co-cluster matrix.
7. The user address consistency assessment method according to claim 1, characterized in that, in, The conflict is determined as follows: when the same user is assigned to different groups in the first user grouping result and the second user grouping result, and the distance between its BeiDou positioning coordinates and the spatial aggregation center of any of its groups exceeds a preset threshold, it is determined that there is a semantic-electrical-spatial consistency conflict. Dynamic anomaly determination is based on tracking the evolution of a user's group within a continuous time window. When a user's trend of leaving the original group or drifting to other groups does not conform to the preset normal evolution pattern learned from historical data, it is determined to be a dynamic evolution anomaly.
8. A user address consistency assessment system, characterized in that, include: The acquisition module is configured to acquire user profile address information, BeiDou positioning time sequence information, and electrical load time sequence data within the target area; The parsing module is configured to perform dual-channel parsing on the file address information, generate multi-level address semantic fingerprint vectors, calculate the address semantic distance between users based on the address semantic fingerprint vectors, and combine preset BeiDou spatial aggregation constraints and preset electrical aggregation constraints to obtain the first user grouping result and the first grouping confidence based on address semantics. The fusion module is configured to calculate the multi-domain electrical distance between users within multiple time windows based on the electrical load time series data, and combine the multi-domain electrical distance, address semantic distance, BeiDou spatial distance and preset structural prior soft constraints, and fuse them through reliability-driven adaptive weights to generate the final distance matrix under spatiotemporal joint evolution. The clustering module is configured to perform clustering based on the final distance matrix to obtain a second user grouping result and a second grouping confidence level based on load characteristics. The determination module is configured to determine that the user address consistency is abnormal when the first user grouping result based on address semantics, the second user grouping result based on load characteristics, and the spatial aggregation result based on BeiDou positioning information conflict or show deviations that do not conform to the preset evolution law within a continuous time window, and output the abnormal evidence chain and abnormal confidence score.
9. An electronic device, characterized in that, include: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method according to any one of claims 1 to 7.