Household transformer abnormity identification method and system based on power utilization address
Through word segmentation analysis of electricity usage addresses and N-gram model to generate address vectors, and calculate the numerical values of abnormal indicators, the problem of inaccurate identification of household variation caused by incomplete latitude and longitude data is solved, and higher recognition accuracy and accurate quantitative evaluation are achieved.
Patent Information
- Application Number
- CN202510850675.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-24
AI Technical Summary
In the prior art, due to incomplete or inaccurate latitude and longitude data, the identification of household variations is inaccurate enough, which affects the efficiency of electricity billing and distribution network management.
Using word segmentation analysis and N-gram model based on electricity usage addresses, address vectors are generated and abnormal indicator values are calculated, and accurate physical address data is used to replace latitude and longitude data for identification.
It improves the accuracy of household variation identification, avoids missed judgments, and realizes accurate quantitative evaluation of the power address of different users.
Smart Images

Figure CN120354849A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of distribution network management, and specifically relates to a method and system for identifying household-transformer anomalies based on electricity consumption addresses. Background Art
[0002] In distribution network management, the relationship between users and substations (transformers) is an important basis for ensuring electricity consumption billing and power supply stability. However, this relationship sometimes exhibits anomalies, which are referred to as household-transformer anomalies. A household-transformer anomaly refers to a phenomenon where the actual correspondence between a user's electricity consumption address and the substation to which the user belongs deviates, for example, the electricity consumption address of a certain user is significantly inconsistent with the distribution of the electricity consumption addresses of other users in the substation. Such anomalies may be caused by factors such as substation division, equipment relocation, data entry errors, etc., and often lead to problems such as inaccurate user billing and decreased distribution network management efficiency. To ensure the stable operation of the power distribution system and the accuracy of electricity consumption data, it is necessary to accurately identify household-transformer anomalies.
[0003] In related technologies, longitude and latitude information is often used to identify household-transformer anomalies. By performing clustering analysis on the geographical coordinates of users' electricity consumption addresses, it is determined whether a user belongs to the normal substation range. However, in practical applications, longitude and latitude data are often difficult to obtain, especially in some areas where address information is incomplete or the geographical information system is not fully covered. Due to the high dependence on longitude and latitude data, when the longitude and latitude data are incomplete or inaccurate, the identification of household-transformer anomalies will be inaccurate. Summary of the Invention
[0004] This application provides a method and system for identifying household-transformer anomalies based on electricity consumption addresses, which can improve the accuracy of identifying household-transformer anomalies.
[0005] In a first aspect, this application provides a method for identifying household-transformer anomalies based on electricity consumption addresses, the method comprising: Obtaining the electricity consumption addresses of multiple users under a substation, and splitting the text of the electricity consumption addresses into address tokens; Counting whether each word in the address vocabulary appears in the corresponding address tokens, and generating an address vector corresponding to the electricity consumption address; Calculating an anomaly index value corresponding to the electricity consumption address according to the address vector; When the anomaly index value meets the anomaly judgment condition, determining the user corresponding to the anomaly index value as an abnormal user.
[0006] By adopting the above technical solution, the electricity consumption address text is segmented and analyzed and converted into a quantifiable address vector, and then the abnormal index value is calculated for abnormal judgment. Compared with the technical solutions in related technologies, accurate physical address data is used to replace the incomplete longitude and latitude data for household transformation abnormality identification, avoiding missed judgments caused by missing parts of the address. At the same time, based on the numerical analysis of the address text and the quantitative judgment of the abnormal index, accurate quantitative evaluation of the electricity consumption addresses of different users is achieved, thereby improving the accuracy of household transformation abnormality identification.
[0007] Optionally, before generating the address vector corresponding to the electricity consumption address by counting whether each word in the statistical address vocabulary appears in the corresponding address segmentation, it further includes: Filter out the unique address segments in the address segmentation and form an initial vocabulary with all the unique address segments; Use the N-gram model to calculate the context occurrence probability corresponding to each word in the initial vocabulary; Sort each word in the initial vocabulary according to the context occurrence probability to obtain the address vocabulary.
[0008] By adopting the above technical solution, the unique address segments in the address segmentation are filtered out and an initial vocabulary is formed, avoiding the interference of repeated address segments on subsequent analysis, enabling the generated address vector to more accurately reflect the characteristic differences of the electricity consumption address, thereby improving the accuracy of household transformation abnormality identification; at the same time, by using the N-gram model to calculate the context occurrence probability between words in the initial vocabulary and sorting the words based on these context occurrence probabilities to generate the address vocabulary, the context relevance and occurrence rules between address words can be effectively captured, enabling the address vocabulary to better reflect the semantic and structural characteristics of the electricity consumption address, thereby improving the accuracy of household transformation abnormality identification.
[0009] Optionally, the using the N-gram model to calculate the context occurrence probability corresponding to each word in the initial vocabulary includes: Add a custom start word to each address segmentation to obtain a training corpus; Build an N-gram model based on the training corpus and use the N-gram model to calculate the context occurrence probability corresponding to each word in the initial vocabulary.
[0010] By adopting the above technical solution, custom starting words are added to the address word segmentation to construct a training corpus, and the N-gram model is used to analyze the context relationship between words, so as to accurately capture the semantic relevance of words in the address text. The N-gram model constructed based on the training corpus can calculate the context occurrence probability of each word in the initial vocabulary, so as to identify the key words with significant features in the address text. This method of lexical analysis based on context probability enables the system to automatically learn and adapt to the characteristics of different address texts, and improves the recognition accuracy of abnormal addresses.
[0011] Optionally, after obtaining the electricity consumption addresses of multiple users in the substation area and splitting the text of the electricity consumption addresses into address word segments, the method further includes: Deleting the head words and tail words in the address word segments.
[0012] By adopting the above technical solution, it is possible to effectively filter out general address words with too high occurrence frequencies (such as provinces, cities, etc.) and atypical address words that have no influence on the result judgment (such as Room 101, Room 102, etc.), so that the subsequent generated address vectors can accurately reflect the key feature information of the electricity consumption addresses, avoiding both the interference caused by general words and the noise and high-dimensional sparsity problems that may be caused by words with no influence, thereby improving the accuracy and discrimination of the address vectors in representing the characteristics of the users' electricity consumption addresses.
[0013] Optionally, the step of counting whether each word in the address vocabulary appears in the corresponding address word segments and generating the address vector corresponding to the electricity consumption address includes: If the word in the address vocabulary appears in the corresponding address word segments, the vector value of the address vector corresponding to the word in the address vocabulary is recorded as 1; If the word in the address vocabulary does not appear in the corresponding address word segments, the vector value of the address vector corresponding to the word in the address vocabulary is recorded as 0.
[0014] By adopting the above technical solution, whether each user's address word segments appear in the address vocabulary is converted into the address vector corresponding to the user in a binary manner, that is, recorded as 1 when it appears and recorded as 0 when it does not appear, which simplifies the expression of address features and reduces the interference effect of the frequency difference of address word segments on abnormal recognition.
[0015] Optionally, the step of calculating the abnormal index value corresponding to the electricity consumption address according to the address vector includes: Obtaining the vector weights corresponding to each vector dimension in the address vector; Substituting the vector weights and the address vector into the abnormal index calculation formula to obtain the abnormal index value corresponding to the electricity consumption address; The calculation formula for the abnormal index is as follows: ; where Z represents the value of the abnormal index, n represents the total number of vector dimensions of the address vector, i represents the i-th vector dimension, represents the vector value corresponding to the i-th vector dimension, represents the vector weight corresponding to the i-th vector dimension.
[0016] By adopting the above technical solution, the vector weights corresponding to each vector dimension in the address vector are obtained, and the vector weights and the address vector are substituted into the calculation formula for the abnormal index to calculate the value of the abnormal index, realizing the differential processing of the importance of different vector dimensions. At the same time, the weighted average method is used to comprehensively evaluate the abnormal degree of the electricity consumption address, so that the value of the abnormal index can more accurately reflect the abnormal characteristics of the electricity consumption address.
[0017] Optionally, when the value of the abnormal index meets the abnormal judgment condition, the user corresponding to the value of the abnormal index is determined as an abnormal user, including: screening out target users whose abnormal index values are less than the first threshold; if the abnormal index values of the remaining users except the target users are all greater than the second threshold, the target users are determined as abnormal users, and the first threshold is less than the second threshold.
[0018] By adopting the above technical solution, target users whose abnormal index values are less than the first threshold are screened out, and it is judged whether the abnormal index values of the remaining users except the target users are all greater than the second threshold to determine abnormal users. The dual-threshold judgment mechanism not only considers the abnormal degree of the target users themselves, but also combines the overall electricity consumption characteristics of the surrounding users for comparative analysis, effectively avoiding misjudgment that may be caused by a single-threshold judgment, thereby improving the accuracy of household-transformer abnormal identification.
[0019] In a second aspect, the present application provides a household-transformer abnormal identification system based on an electricity consumption address. The system includes: An initialization module, configured to obtain the electricity consumption addresses of multiple users in a substation area and split the text of the electricity consumption addresses into address segments; A vector conversion module, configured to count whether each word in the address vocabulary appears in the corresponding address segments and generate an address vector corresponding to the electricity consumption address; A calculation module, configured to calculate the value of the abnormal index corresponding to the electricity consumption address according to the address vector; An identification module, configured to determine the user corresponding to the value of the abnormal index as an abnormal user when the value of the abnormal index meets the abnormal judgment condition.
[0020] In a third aspect, the present application provides a computer storage medium storing a plurality of instructions adapted to be loaded and executed by a processor to perform any of the above methods.
[0021] In a fourth aspect, the present application provides an electronic device including a processor, a memory, and a transceiver. The memory is used to store instructions, the transceiver is used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device performs any of the above methods.
[0022] In summary, the beneficial effects brought by the technical solution of the present application include: By adopting the above technical solution, the electricity consumption address text is segmented and analyzed and converted into a quantifiable address vector, and then the abnormal index value is calculated for abnormal judgment. Compared with the technical solutions in the related art, accurate physical address data is used to replace the incomplete longitude and latitude data for household change abnormality identification, avoiding missed judgments caused by missing parts of the address. At the same time, based on the numerical analysis of the address text and the quantitative judgment of the abnormal index, accurate quantitative evaluation of the electricity consumption addresses of different users is realized, thereby improving the accuracy of household change abnormality identification. Description of the Drawings
[0023] Figure 1 is a schematic flowchart of a household change abnormality identification method based on electricity consumption address according to an embodiment of the present application; Figure 2 is a schematic structural diagram of a household change abnormality identification system based on electricity consumption address according to an embodiment of the present application; Figure 3 is a schematic structural diagram of an electronic device provided by an embodiment of the present application.
[0024] Description of the reference numerals: 300, electronic device; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. Detailed Embodiments
[0025] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments.
[0026] In the description of the embodiments of this application, words such as "exemplary", "for example", or "for instance" are used to give examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary", "for example", or "for instance" in the embodiments of this application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary", "for example", or "for instance" is intended to present relevant concepts in a specific manner.
[0027] In the description of the embodiments of this application, the term "plurality" means two or more. For example, a plurality of systems means two or more systems, and a plurality of screen terminals means two or more screen terminals. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "comprise", "include", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0028] Please refer to Figure 1 , which is a schematic flowchart of a method for identifying household-transformer anomalies based on electricity consumption addresses provided by the embodiments of this application. This method can be implemented depending on a computer program, can be implemented depending on a single-chip microcomputer, or can run on a system for identifying household-transformer anomalies based on electricity consumption addresses based on the von Neumann architecture. This computer program can be integrated into an application or can run as an independent tool-class application. The following will elaborate on the specific steps of the method for identifying household-transformer anomalies based on electricity consumption addresses.
[0029] S101: Obtain the electricity consumption addresses of multiple users in a substation area, and split the text of the electricity consumption addresses into address segments; Among them, the electricity consumption address refers to the detailed geographical location information where the user's electricity consumption facilities are located. In the embodiments of this application, it can be understood as a standardized text address including hierarchical structures such as province, city, district (county), street (township), community (village), building number, unit number, and room number, which is used to uniquely identify the power supply service location of each electricity consumption user and serves as an important basis for user identity recognition. The following gives Table 1 to understand the electricity consumption address.
[0030] Among them, the address segment refers to the smallest semantic unit obtained by splitting the complete text of the electricity consumption address according to semantics and address composition elements. In the embodiments of this application, it can be understood as the independent address tokens obtained by dividing the standardized text address including hierarchical structures such as province, city, district (county), street (township), community (village), building number, unit number, and room number, which is used to realize the normalization processing and vector representation of the address text.
[0031] In specific implementation, first, obtain the electricity consumption address information of all users in the substation area from the power system database. Since the electricity consumption address is a standardized text address with a hierarchical structure including province, city, district (county), street (township), community (village), building number, unit number, room number, etc., these addresses are usually stored in the form of continuous strings, which is not convenient for direct anomaly analysis. Therefore, it is necessary to use Chinese word segmentation technology to split the obtained electricity consumption address text, and split each electricity consumption address text into multiple address words with independent semantics. Specifically, the jieba word segmentation algorithm can be used, combined with the address dictionary unique to the power industry, to segment the electricity consumption address text.
[0032] To facilitate understanding of the process of segmenting electricity consumption addresses, please refer to Table 1 and Table 2 specifically. Table 1 shows the original electricity consumption addresses of 7 users in the substation area. These addresses are all located in D Street, C District, B City, A Province, but are distributed in different communities, buildings and room numbers. Table 2 shows the results after segmenting these addresses. Each address is split into multiple independent semantic units, including the province "A Province", the city "B City", the region "C District", the street "D Street", the community names "Community A / Community B", the building numbers "Building 2 / Building 3" and the room numbers. This word segmentation method can convert the continuous address text into an address word sequence.
[0033] Table 1 - Example Table of Electricity Consumption Addresses:
[0034] Table 2 - Example Table of Address Word Segmentation Results:
[0035] Based on the above embodiments, as an optional implementation manner, the steps of generating an address vocabulary specifically include S201 - S203.
[0036] S201: Screen out the unique address words in the address words, and form an initial vocabulary with all the unique address words.
[0037] Specifically, before performing electricity consumption address anomaly analysis, it is necessary to sort out and remove duplicates from the address words of all users in the substation area. Since each user's electricity consumption address is segmented into multiple address words, and these address words are repeated among different users, such as the same administrative division information of province, city, district and county, etc., it is necessary to traverse all users' address words and use a set data structure to remove duplicate address words to obtain an initial vocabulary containing all unique address words. For example, the initial vocabulary that can be formed by the above Table 2 is: 101, 102, 2, 201, 202, 3, Community B, D Street, Building, B City, A Province, Community A, C District.
[0038] S202: Calculate the context occurrence probability corresponding to each vocabulary in the initial vocabulary using the N-gram model; Among them, the N-gram model refers to a language model that segments text in the way of consecutive N characters or words. In the embodiments of this application, it can be understood that the vocabularies in the address vocabulary are divided into several overlapping sub-string sequences according to consecutive N characters, so as to capture the local consecutive features of the characters in the table.
[0039] Among them, the address vocabulary refers to the initial vocabulary after sorting.
[0040] In specific implementation, the N-gram model can be used to predict the probability distribution of the vocabulary sequence in the text. Its core assumption is that the occurrence probability of the Nth word is only related to the previous N-1 words, rather than the entire context. The N-gram model is a commonly used statistical language model in natural language processing (NLP), which is used to predict the next element in the sequence (such as the next word in the text). The N-gram model can be used to sort the vocabularies in the bag-of-words model vocabulary, so that the vocabularies with strong local relevance are arranged adjacent to each other, thereby obtaining an address vocabulary arranged from large to small by region.
[0041] In specific implementation, first add a custom start word (such as "START") at the start position of each address word segmentation to mark the start boundary of the address text, thereby constructing the training corpus. For example, for the address word segmentation "a certain community a certain street", it becomes "START a certain community a certain street" after adding the start word. Then, based on these training corpora, construct an N-gram model. When N takes the value of 2 (i.e., the Bi-gram model), the system will count the occurrence frequency of the word pairs and calculate the conditional probability of the next word given the previous word. For example, the probability of "street" appearing after "community" can be calculated, so as to obtain the context occurrence probability of each vocabulary in the initial vocabulary. This probability-based analysis method can capture the correlation between the vocabularies in the address text and effectively identify the key vocabularies in the address structure. The context occurrence probability calculated not only reflects the importance of the vocabulary in the address text, but also contains the position information and semantic correlation information of the vocabulary. These probability values provide a quantitative basis for subsequent vocabulary sorting and screening, enabling the system to construct an address vocabulary containing significant feature vocabularies, thereby improving the expression ability of the address vector and the accuracy of anomaly recognition. At the same time, this N-gram-based probability calculation method has the characteristics of strong interpretability and high calculation efficiency, and is suitable for processing large-scale electricity consumption address data.
[0042] For the step of calculating the context occurrence probability corresponding to each word in the initial vocabulary using the N-gram model, a specific implementation is as follows: Add a custom start word to each of the address segmentations to obtain a training corpus; build an N-gram model based on the training corpus, and use the N-gram model to calculate the context occurrence probability corresponding to each word in the initial vocabulary.
[0043] In specific implementation, first add a custom start word "START" at the start position of the address segmentation sequence of each user. For example, process the address segmentation "Area C, Street D" as "START Area C, Street D". All the address segmentation sequences processed in this way constitute the training corpus. The purpose of adding the start word is to mark the start boundary of each address text, enabling the system to recognize and learn the characteristics of the appearance of words at the start position of the address. Then, build an N-gram model based on these training corpora. When using the bigram model, the system will calculate the conditional probability P(word | previous word), especially P(word | "START"), that is, calculate the probability of a certain word appearing as the first segmentation of the address. For example, the system can calculate the occurrence probability of the word "Area" after "START" and before "C", so as to obtain the probability distribution of this word at different context positions. This N-gram-based probability calculation method not only considers the independent occurrence frequency of words, but also includes the positional association information between words, enabling the system to accurately grasp the occurrence rules of words in the address text.
[0044] S203: Sort each word in the initial vocabulary according to the context occurrence probability to obtain an address vocabulary.
[0045] In specific implementation, the system uses the context occurrence probability calculated by the N-gram model as the sorting basis. For example, for the calculated conditional probabilities P(word | "START") and P(word | previous word), the system will sort the words in the initial vocabulary in descending order according to these probability values. A high probability value indicates that the word is more likely to appear in a specific context and has stronger address characteristics.
[0046] Based on the above embodiments, as an optional implementation, delete the head words and tail words in the address segmentation.
[0047] To address the problem of high-dimensional sparsity caused by a large vocabulary in the address vector, a method of processing the head and tail words in the address word segmentation is adopted to reduce the vector dimension. In actual electricity consumption scenarios, normal users under the same transformer area often have similar addresses, and the main differences are reflected in the end information of the address such as the house number, such as the difference in room numbers like "101" and "102". Since this end difference does not affect the judgment of household-transformer abnormality, the words that only appear at the end can be deleted. At the same time, considering that household-transformer abnormality usually does not cross urban areas, the administrative division information at the head of the address, such as words like "Province A" and "City B", contributes less to abnormality recognition and can also be deleted. Through this targeted dimension reduction method, not only the address features that are most crucial for identifying household-transformer abnormalities are retained, but also the vector dimension is significantly reduced, improving the computational efficiency of the algorithm.
[0048] Exemplarily, the address vector obtained after processing the address word segmentation result example table in Table 2 through the above steps is shown in Table 3, and the word order in Table 3 has been sorted through the above steps.
[0049] Table 3 - Address Vector Result Table:
[0050] S102: Count whether each word in the address vocabulary appears in the corresponding address word segmentation, and generate an address vector corresponding to the electricity consumption address.
[0051] This application adopts the Bag of Words (BoW) model to convert text data into numerical data. The Bag of Words model is a model in natural language processing and text mining. Its principle is to establish a vocabulary that contains all the unique words that appear in the text, and then for each text, count whether each word in the vocabulary appears in the text to form a vector of the same length as the vocabulary.
[0052] Among them, the address vector refers to converting the address word segmentation of the user's electricity consumption address into a numerical form. In the embodiments of this application, it can be understood that based on the Bag of Words model, after establishing a vocabulary for the address word segmentation of all users in the transformer area, count whether each word in the address vocabulary appears in the corresponding address word segmentation to obtain a numerical vector, which is used to convert the electricity consumption address in text form into a numerical representation that can perform mathematical operations.
[0053] Among them, the address vocabulary refers to a non-repeating ordered vocabulary set formed by segmenting the electricity consumption addresses of all users within the substation area. In the embodiments of this application, it can be understood as a list containing all unique address segments such as province, city, district, county, street, community, building number, unit number, room number, etc., obtained by removing duplicates from the address segments of all users within the substation area, and is used to construct a bag-of-words model to convert the electricity consumption address into an address vector, realizing the mapping from address text to numerical form.
[0054] For each electricity consumption address, the system matches its address segments with the vocabulary in the address vocabulary. When a vocabulary in the address vocabulary appears in the corresponding address segment, the vector value corresponding to the vocabulary in the address vector is marked as 1, otherwise it is marked as 0. For example, if the address segments of A are 1, 3, 5 and the address vocabulary is 1, 2, 3, 4, 5, 6, then the address vector of A is 1, 0, 1, 0, 1, 0.
[0055] Specifically, step S102 specifically includes the following steps: If the vocabulary in the address vocabulary appears in the corresponding address segment, then record the vector value of the address vector corresponding to the vocabulary in the address vocabulary as 1; if the vocabulary in the address vocabulary does not appear in the corresponding address segment, then record the vector value of the address vector corresponding to the vocabulary in the address vocabulary as 0.
[0056] When constructing the term frequency vector, considering the particularity of the electricity consumption address, the same address segment usually only appears once in the electricity consumption address of a single user. For example, an address will not repeatedly appear as "Province A" or multiple room numbers. Therefore, it is more in line with the actual situation to construct the term frequency vector in a binary way. Specifically, when implementing, for the electricity consumption address of each user, traverse the vocabulary in the address vocabulary and check one by one whether each vocabulary appears in the corresponding address segment: If the vocabulary in the address vocabulary appears in the corresponding address segment, then record the vector value of the address vector corresponding to the vocabulary in the address vocabulary as 1; if the vocabulary in the address vocabulary does not appear in the corresponding address segment, then record the vector value of the address vector corresponding to the vocabulary in the address vocabulary as 0.
[0057] S103: Calculate the abnormal index value corresponding to the electricity consumption address according to the address vector; Among them, the abnormal index value refers to a quantitative index for measuring the degree of difference between the electricity consumption address and the addresses of other users within its affiliated substation area. In the embodiments of this application, it can be understood as a value obtained by calculating based on the address vector and representing the abnormal degree of a single electricity consumption address, and is used to quantitatively evaluate whether the electricity consumption address is an abnormal address of the household-transformer.
[0058] Specifically, for each address vector, a set of difference values is obtained by calculating the difference features between it and the address vectors of other users in the same power distribution area. Then, an anomaly index is constructed by statistically analyzing the distribution features of this set of difference values. This method for calculating the anomaly index based on vector differences can effectively measure the degree of difference between a single electricity consumption address and other addresses in its affiliated power distribution area. The greater the difference, the greater the difference between this address and the addresses of other users in the power distribution area, and the more likely it is to be a user with abnormal household-transformer connection.
[0059] Based on the above embodiments, as an alternative implementation manner, the anomaly index value can be calculated as follows: Obtain the vector weights corresponding to each vector dimension in the address vector; Substitute the vector weights and the address vector into the anomaly index calculation formula to obtain the anomaly index value corresponding to the electricity consumption address; The anomaly index calculation formula is: ; where Z represents the anomaly index value, n represents the total number of vector dimensions of the address vector, i represents the i-th vector dimension, represents the vector value corresponding to the i-th vector dimension, represents the vector weight corresponding to the i-th vector dimension.
[0060] First, obtain the vector weights corresponding to each vector dimension in the address vector, and these weights reflect the importance of each dimension in identifying abnormal household-transformer connections. Then, substitute the obtained vector weights and the address vector into the anomaly index calculation formula, where the anomaly index value is equal to the sum of the products of the vector values corresponding to the vector dimensions and the vector weights divided by the sum of the vector weights. The purpose of adopting this weighted calculation method is to take into account that different dimensions in the address vector contribute differently to the judgment of abnormal household-transformer connections. By introducing vector weights, the influence of important dimensions can be highlighted and the influence of secondary dimensions can be weakened. For example, a higher weight is assigned to the dimension that can better reflect the core features of the address, while a lower weight is assigned to the dimension of detailed information, so that the calculated anomaly index value can more accurately reflect the degree of anomaly of the electricity consumption address.
[0061] Regarding the vector weight, the vector weight is a quantitative representation of the importance of each dimension of the address vector, and its setting needs to comprehensively consider the actual business scenario and data characteristics. In practical applications, the vector weight can be determined from the following aspects: First, since the address vocabulary has been sorted before, the more forward the different words appear, the greater the address difference. For example, the difference at the "street" level is greater than that at the "community" level because the more forward the word, the greater the weight. Second, based on the household change anomaly cases found in historical data, the contribution of different address dimensions to anomaly recognition can be analyzed, and higher weights can be assigned to the dimensions with good recognition effects. When implementing specifically, a data-driven approach can be adopted. By analyzing a large amount of historical data, the discrimination ability of each dimension in distinguishing normal users from abnormal users can be calculated, and this discrimination ability can be converted into corresponding weight values.
[0062] Exemplarily, the present application calculates the address vector in Table 3 and obtains the distribution table of anomaly index values shown in Table 4.
[0063] Table 4 - Distribution Table of Anomaly Index Values:
[0064] S104: When the anomaly index value meets the anomaly judgment condition, the user corresponding to the anomaly index value is determined as an abnormal user.
[0065] In this embodiment, by comparing the anomaly index value with the preset anomaly judgment condition, abnormal users in household change can be effectively identified. When the calculated anomaly index value meets the anomaly judgment condition, the user corresponding to this anomaly index value can be determined as an abnormal user.
[0066] Based on the above embodiment, as an optional implementation manner, step S104 specifically further includes steps S301 - S302.
[0067] The anomaly judgment condition is divided into two judgment steps to accurately identify abnormal users.
[0068] S301: Screen out the target users whose anomaly index values are less than the first threshold; In this embodiment, by setting a first threshold and screening target users whose abnormal index values are less than this threshold, potential abnormal users in terms of household changes can be initially identified. Specifically in implementation, first, a suitable first threshold is set according to historical data and actual business experience, and then the abnormal index values of all users are compared with this threshold, and users with abnormal index values less than the first threshold are screened out as target users. The purpose of setting the first threshold is to establish a preliminary screening criterion. Because the smaller the abnormal index value, it indicates that the electricity consumption address of this user has a greater difference from the addresses of other users in the substation area, and there is a greater possibility of abnormal household changes. This screening method based on the threshold can quickly locate the group of suspicious users and improve the efficiency of subsequent verification work.
[0069] S302: If the abnormal index values of the remaining users except the target users are all greater than the second threshold, then determine the target user as an abnormal user, and the first threshold is less than the second threshold.
[0070] By introducing the second threshold and combining it with the first threshold to construct a dual-threshold judgment mechanism, abnormal users in terms of household changes can be identified more accurately. Specifically in implementation, first ensure that the set first threshold is less than the second threshold, and then check whether the abnormal index values of the remaining users except the target user are all greater than the second threshold. If this condition is met, then determine the target user as an abnormal user. The purpose of adopting this dual-threshold judgment mechanism is to further verify the abnormality of the target user by comparing and analyzing the distribution of abnormal index values between the target user and other users. When the abnormal index value of the target user is significantly lower than the first threshold, while the abnormal index values of other users are all higher than the second threshold, it indicates that the address characteristics of this target user have a significant difference from those of the surrounding users, and the degree of this difference exceeds the normal fluctuation range. Therefore, it can be more reliably determined that it is an abnormal user in terms of household changes. This judgment method based on dual thresholds can effectively reduce the misjudgment rate and improve the accuracy of identifying abnormal users in terms of household changes. For example, when there is an obvious break in the distribution of abnormal index values of a certain user in the substation area (such as User Seven in Table 3), that is, the abnormal index value of an individual user is significantly lower than the first threshold, while the abnormal index values of other users are all higher than the second threshold, the real abnormal users in terms of household changes can be more accurately identified.
[0071] The following is an embodiment of the system of this application, which can be used to execute the embodiment of the method of this application. For details not disclosed in the embodiment of the system of this application, please refer to the embodiment of the method of this application.
[0072] Please refer to Figure 2 , which shows a schematic structural diagram of a system for identifying abnormal household changes based on electricity consumption address provided by an exemplary embodiment of this application. This system can be implemented as all or part of the system through a combination of software, hardware, or both. The system for identifying abnormal household changes based on electricity consumption address includes: An initialization module, configured to obtain the electricity consumption addresses of multiple users in a power distribution area, and split the text of the electricity consumption addresses into address word segments; A vector conversion module, configured to count whether each word in an address vocabulary appears in the corresponding address word segment, and generate an address vector corresponding to the electricity consumption address; A calculation module, configured to calculate an abnormal index value corresponding to the electricity consumption address according to the address vector; An identification module, configured to, when the abnormal index value meets an abnormal judgment condition, determine the user corresponding to the abnormal index value as an abnormal user.
[0073] Based on the above embodiments, as an alternative embodiment, the initialization module is further configured to filter out unique address word segments in the address word segments, and form all the unique address word segments into an initial vocabulary; use an N-gram model to calculate the context occurrence probability corresponding to each word in the initial vocabulary; sort each word in the initial vocabulary according to the context occurrence probability to obtain an address vocabulary.
[0074] Based on the above embodiments, as an alternative embodiment, the initialization module is further configured to add a custom start word to each of the address word segments to obtain a training corpus; construct an N-gram model based on the training corpus, and use the N-gram model to calculate the context occurrence probability corresponding to each word in the initial vocabulary.
[0075] Based on the above embodiments, as an alternative embodiment, the initialization module is further configured to delete the head words and tail words in the address word segments.
[0076] Based on the above embodiments, as an alternative embodiment, the vector conversion module is further configured to, if a word in the address vocabulary appears in the corresponding address word segment, set the vector value of the address vector corresponding to the word in the address vocabulary to 1; if a word in the address vocabulary does not appear in the corresponding address word segment, set the vector value of the address vector corresponding to the word in the address vocabulary to 0.
[0077] Based on the above embodiments, as an alternative embodiment, the calculation module is further configured to obtain vector weights corresponding to each vector dimension in the address vector; substitute the vector weights and the address vector into an abnormal index calculation formula to obtain an abnormal index value corresponding to the electricity consumption address; the abnormal index calculation formula is: ; where Z represents the abnormal index value, n represents the total number of vector dimensions of the address vector, i represents the i-th vector dimension, represents the vector value corresponding to the i-th vector dimension, represents the vector weight corresponding to the i-th vector dimension.
[0078] Based on the above embodiments, as an alternative embodiment, the recognition module is further configured to screen out target users whose abnormal index values are less than a first threshold; if the abnormal index values of the remaining users except the target users are all greater than a second threshold, the target users are determined as abnormal users, and the first threshold is less than the second threshold.
[0079] The embodiment of the present application also provides a computer storage medium, which can store multiple instructions. The instructions are suitable for being loaded and executed by a processor to perform the household- transformer abnormality recognition method based on the electricity consumption address as described in the above embodiments. The specific execution process can refer to the specific description of the embodiments and will not be elaborated here.
[0080] Please refer to Figure 3 , which is a schematic structural diagram of an electronic device provided by the embodiment of the present application. As Figure 3 shown, the electronic device 300 may include: at least one processor 301, at least one network interface 304, a user interface 303, a memory 305, and at least one communication bus 302.
[0081] Among them, the communication bus 302 is used to implement connection communication between these components.
[0082] Among them, the user interface 303 may include a standard wired interface and a wireless interface.
[0083] Among them, the network interface 304 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).
[0084] Among them, the processor 301 may include one or more processing cores. The processor 301 connects various parts within the entire server through various interfaces and lines. By running or executing instructions, programs, code sets, or instruction sets stored in the memory 305, and by calling the data stored in the memory 305, it executes various functions of the server and processes data. Optionally, the processor 301 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 301 may integrate one or a combination of several of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor 301 and may be implemented separately by a single chip.
[0085] Among them, the memory 305 may include random access memory (RAM) and may also include read-only memory. Optionally, the memory 305 includes a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 305 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store the data involved in the above-mentioned various method embodiments. Optionally, the memory 305 may also be at least one storage device located far from the aforementioned processor 301. As Figure 3 shown, the memory 305, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for a household transformation anomaly recognition method based on electricity consumption addresses.
[0086] In Figure 3In the electronic device 300 shown, the user interface 303 is mainly used to provide an interface for the user to input and obtain the data input by the user; and the processor 301 can be used to call an application program stored in the memory 305, which is a method for identifying household transformer anomalies based on electricity consumption addresses. When executed by one or more processors, the electronic device executes one or more of the methods in the above embodiments.
[0087] An electronic device-readable storage medium stores instructions. When executed by one or more processors, the electronic device executes one or more of the methods in the above embodiments.
[0088] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0089] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0090] In several embodiments provided by this application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some service interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.
[0091] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0092] In addition, in each embodiment of this application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0093] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned memory includes various media that can store program codes, such as USB flash drives, mobile hard disks, magnetic disks, or optical discs.
[0094] The above are only exemplary embodiments of the present disclosure, and the scope of the present disclosure cannot be limited thereby. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. After considering the specification and the disclosure of the practical truth, those skilled in the art will easily think of other implementation manners of the present disclosure. The present application aims to cover any variations, uses, or adaptive changes of the present disclosure, and these variations, uses, or adaptive changes follow the general principles of the present disclosure and include common general knowledge or conventional technical means in the technical field not recorded in the present disclosure.
Claims
1. A method for identifying household transformer anomalies based on electricity consumption addresses, characterized in that, The method includes: Obtaining the electricity consumption addresses of multiple users in the substation area, and splitting the text of the electricity consumption addresses into address segmentation words; Counting whether each word in the address vocabulary appears in the corresponding address segmentation words, and generating an address vector corresponding to the electricity consumption address; Calculating an abnormal index value corresponding to the electricity consumption address according to the address vector; When the abnormal index value meets the abnormal judgment condition, determining the user corresponding to the abnormal index value as an abnormal user.
2. The method according to claim 1, wherein Before counting whether each word in the address vocabulary appears in the corresponding address segmentation words and generating an address vector corresponding to the electricity consumption address, it further includes: Filtering out the unique address segmentation words in the address segmentation words, and forming an initial vocabulary with all the unique address segmentation words; Using an N-gram model to calculate the context occurrence probability corresponding to each word in the initial vocabulary; Sorting each word in the initial vocabulary according to the context occurrence probability to obtain an address vocabulary.
3. The method according to claim 2, wherein The using an N-gram model to calculate the context occurrence probability corresponding to each word in the initial vocabulary includes: Adding a custom start word to each of the address segmentation words to obtain a training corpus; Constructing an N-gram model based on the training corpus, and using the N-gram model to calculate the context occurrence probability corresponding to each word in the initial vocabulary.
4. The method according to claim 1, characterized in that, After obtaining the electricity consumption addresses of multiple users in the substation area and splitting the text of the electricity consumption addresses into address segmentation words, it further includes: Deleting the head words and tail words in the address segmentation words.
5. The method according to claim 1, characterized in that The counting whether each word in the address vocabulary appears in the corresponding address segmentation words and generating an address vector corresponding to the electricity consumption address includes: If the word in the address vocabulary appears in the corresponding address segmentation words, setting the vector value of the address vector corresponding to the word in the address vocabulary to 1; If the word in the address vocabulary does not appear in the corresponding address segmentation words, setting the vector value of the address vector corresponding to the word in the address vocabulary to 0.
6. The method according to claim 1, wherein The calculating an abnormal index value corresponding to the electricity consumption address according to the address vector includes: Obtaining the vector weights corresponding to each vector dimension in the address vector; Substituting the vector weights and the address vector into an abnormal index calculation formula to obtain an abnormal index value corresponding to the electricity consumption address; The calculation formula for the abnormal index is as follows: ; Where Z represents the abnormal index value, n represents the total number of vector dimensions of the address vector, and i represents the i-th vector dimension, represents the vector value corresponding to the i-th vector dimension, represents the vector weight corresponding to the i-th vector dimension.
7. The method according to claim 1, characterized in that, The when the abnormal index value meets the abnormal judgment condition, determining the user corresponding to the abnormal index value as an abnormal user includes: Filtering out target users with abnormal index values less than a first threshold; If the abnormal index values of the remaining users except the target users are all greater than the second threshold, determining the target users as abnormal users, where the first threshold is less than the second threshold.
8. An abnormal identification system for household-transformer based on electricity consumption address, characterized in that, The system includes: An initialization module, configured to obtain the electricity consumption addresses of multiple users in the substation area, and split the text of the electricity consumption addresses into address segmentation words; A vector conversion module, configured to count whether each word in the address vocabulary appears in the corresponding address segmentation words, and generate an address vector corresponding to the electricity consumption address; A calculation module, configured to calculate an abnormal index value corresponding to the electricity consumption address according to the address vector; An identification module, configured to, when the abnormal index value meets an abnormal judgment condition, determine a user corresponding to the abnormal index value as an abnormal user.
9. A computer storage medium, characterized in that, The computer storage medium stores a plurality of instructions, and the instructions are adapted to be loaded and executed by a processor to perform the method according to any one of claims 1 to 7.
10. An electronic device, characterized in that, It includes a processor, a memory, and a transceiver. The memory is used to store instructions, the transceiver is used to communicate with other devices, and the processor is used to execute the instructions stored in the memory so that the electronic device performs the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method and device for retrieving work order address of power distribution network
CN106503033A
Recognition method for abnormal user of power system
CN107958395A
A method and system for identifying household transformer relationship in a low-pressure platform area
CN109344144A
A method for judging the abnormality of user change relation based on the user address
CN109447490A
Abnormality detection method, device and equipment in transformer area user transformer cutover scene and storage medium
CN119337269A