Customer data processing method and system
By constructing a state transition diagram and a collaborative state mapping matrix, the dynamic modeling problem of identifying abnormal behavior of cross-regional power customers was solved, enabling accurate identification and dynamic tracing of abnormal behavior, improving identification accuracy and modeling depth, and possessing continuous learning and adaptive optimization capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- STATE GRID ZHEJIANG ELECTRIC POWER CO MARKETING SERVICE CENT
- Filing Date
- 2025-08-01
- Publication Date
- 2026-04-10
AI Technical Summary
Existing technologies struggle to effectively reflect the dynamic migration trends of customer response strategies when dealing with anomaly identification and structural modeling of regional distributed power customer behavior. Furthermore, they lack the ability to conduct cross-regional collaborative anomaly detection and dynamic correction of state diagram structures, making it difficult to share and compare behavioral characteristics across regions and resulting in insufficient identification capabilities.
A state transition graph is constructed, atypical state sequences are extracted through a bidirectional state structure, a collaborative state mapping matrix is constructed, the correlation of abnormal transitions is quantified, distributed abnormal state fragments are identified, and a structured feature set is generated through temporal coding. The node clustering boundaries and transition edge weights of the state transition graph are corrected to form a dynamic evolutionary expression structure.
It achieves accurate identification and dynamic tracing of abnormal behavior, improves the identification accuracy of cross-regional synchronous abnormal patterns, enhances the expressive power of data modeling and the real-time performance of system response, and has the ability to continuously learn and adaptively optimize.
Smart Images

Figure CN120912244B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data processing, in particular to a customer data processing method and system. BACKGROUND
[0002] Under the background of the increasing intelligence of modern power systems, the collection and processing of customer-side power consumption behavior data have become the key foundation for promoting the upgrading of precise load forecasting, anomaly detection and intelligent scheduling and other business capabilities. With the widespread deployment of smart meters and multi-source sensing terminals, power companies can obtain multi-dimensional power consumption event sequences containing customer power consumption power, time period characteristics, device start-stop events and the like. Based on these event sequences, researchers have proposed various behavior modeling and graph structure analysis methods, such as state transition graph construction, Markov models, sequence clustering and graph neural networks, for mining and predicting customer behavior characteristics.
[0003] However, due to significant regional differences, diverse customer behavior patterns, and complex evolution trends of power usage scenarios, existing methods are often limited to static structure modeling or behavior clustering within a single region, making it difficult to meet the needs of efficient perception and evolution modeling of dynamic, cross-regional behavior anomalies.
[0004] The existing technology has multiple bottlenecks to be broken through in the aspect of processing regional distributed power customer behavior anomaly identification and structure modeling. On the one hand, typical methods mostly rely on statistical characteristics or frequent behavior segments within a single region, ignoring the temporal coupling relationship in the evolution process of customer behavior, especially in the identification of low-frequency atypical event sequences, which easily leads to the omission of key behavior segments. On the other hand, in multi-regional data fusion and collaborative analysis, existing methods lack systematic state structure mapping and cross-regional anomaly association modeling mechanisms, making it difficult to share and compare behavior characteristics between regions, hindering the identification and expression of large-scale power customer behavior anomaly aggregation effects. In addition, existing state transition modeling mostly uses fixed edge weights or rule updating methods, lacking sensitivity to dynamic changes in events, making it difficult to effectively reflect the dynamic migration trends of customer response strategies in the time evolution process. SUMMARY
[0005] In view of the problems of existing customer data processing technology in atypical state sequence mining, cross-regional collaborative anomaly detection and state graph structure dynamic correction, the present application is proposed.
[0006] Therefore, the problem to be solved by the present application is how to improve the identification accuracy of abnormal behavior segments and the expression depth of behavior response modeling.
[0007] To solve the above technical problems, the present application provides the following technical solutions:
[0008] In a first aspect, the present invention provides a customer data processing method, comprising: initially cleaning the electricity consumption event sequences of customers in various regional power grids and constructing a state transition diagram; constructing a bidirectional state structure for the state transition diagram, extracting atypical state sequences based on the path with the maximum transition probability in each region, and labeling them as task key frame state sequences within the region; constructing a collaborative state mapping matrix based on the task key frame state sequences extracted from each region, quantifying the correlation of abnormal transitions between different regions through behavioral difference vectors, and identifying distributed state anomaly segments with correlation and aggregation characteristics; performing time-series encoding on the task key frame state sequences corresponding to the state anomaly segments with correlation and aggregation characteristics in the collaborative state mapping matrix to generate a structured feature set representing the phased response behavior of customers; and correcting the clustering boundaries and transition edge weights of state nodes in the state transition diagram based on the structured feature set to form a dynamic evolutionary expression structure oriented towards behavioral anomalies in a specific region.
[0009] As a preferred embodiment of the customer data processing method of the present invention, the state nodes in the state transition diagram are determined based on the clustering of customers' historical electricity consumption patterns, and the weights of the transition edges are calculated based on the electricity consumption frequency and the degree of behavioral mutation.
[0010] As a preferred embodiment of the customer data processing method of the present invention, the construction of the state transition diagram includes: using an embedded spectral clustering method to compress the state space of the cleaned electricity consumption event segments and map the event groups to state nodes; after the state nodes are determined, calculating the state transition probability based on the time sequence order and mutation weight of adjacent states within the electricity consumption event segments, and constructing the state transition diagram.
[0011] As a preferred embodiment of the customer data processing method of the present invention, the construction of the bidirectional state structure includes: for state node pairs with transition relationships in the state transition graph, calculating the forward and reverse transition probabilities respectively; wherein, the forward transition probability is the state transition probability of the starting state node jumping to the target state node; the reverse transition probability is calculated by taking the target state node as the backtracking starting point, extracting the reverse jump path and calculating the corresponding reverse transition frequency, and finally constructing a bidirectional symmetric state transition probability matrix.
[0012] As a preferred embodiment of the customer data processing method of the present invention, the reverse transition frequency is calculated as follows: starting from each target state node, traversing backward along the time sequence of the state transition graph, limiting the backtracking step window and time span threshold, extracting possible preceding state path segments, and recording the reverse transition frequency between state pairs; performing frequency statistics on all state pairs in the extracted reverse state paths, and combining the mutation weight, normalizing to obtain the reverse transition probability value of each state pair.
[0013] As a preferred scheme of the customer data processing method, the extraction of the atypical state sequence comprises: for each regional state transition graph, constructing a path cumulative probability based on state transition probability products, and filtering out an atypical path sequence with the maximum cumulative transition probability value and a frequency lower than a historical average value from each state node to a corresponding periodic terminal state node by using a probability maximum path extraction strategy.
[0014] As a preferred scheme of the customer data processing method, the construction of the cooperative state mapping matrix comprises: for a task key frame state sequence in each region, constructing a behavior feature vector sequence based on state jump amplitudes and time position differences; wherein the jump amplitude is calculated according to a path transition weight difference degree before and after the key frame; performing bidirectional mapping comparison on the behavior feature vector sequences of different regions to construct a behavior similarity matrix, and retaining a high-similarity state pair with a similarity higher than a similarity threshold; on the basis of the behavior similarity matrix, the number of times that the high-similarity state pair appears together in all regional combinations is counted; when the frequency of the high-similarity state pair appearing together in the regional combinations exceeds a support threshold, the high-similarity state pair is determined as a cross-regional cooperative state pair, and a cooperative state mapping matrix is constructed, and a regional combination index and a behavior difference index vector are marked.
[0015] As a preferred scheme of the customer data processing method, the identification of the distributed state abnormal segment with the correlation aggregation characteristic comprises: according to the behavior difference index vector in the cooperative state mapping matrix, a state segment with a similar transition structure in multiple regions is divided as the distributed state abnormal segment with the correlation aggregation characteristic.
[0016] As a preferred scheme of the customer data processing method, the division of the state segment with the similar transition structure in multiple regions comprises: filtering out continuous high-similarity state pairs from the cooperative state mapping matrix to combine into multiple state segments; comparing and analyzing the behavior difference indexes involved in each state segment to determine whether the performances in different regions are similar; when the state segment shows a similar transition structure in multiple regions and the behavior difference indexes meet a set threshold, the corresponding state segment is determined as the distributed state abnormal segment with the correlation aggregation characteristic.
[0017] As a preferred scheme of the customer data processing method, the time sequence coding comprises: based on the distributed state abnormal segment, a corresponding multiple-regional task key frame state sequence set is extracted from the cooperative state mapping matrix to construct a cross-regional time sequence transition window, and the multiple-regional task key frame state sequence set is uniformly coded as a multiple-regional linkage state segment; for the multiple-regional linkage state segment, a multi-dimensional time sequence coding method is used to fuse the transition probability, jump gradient and regional difference features of the state nodes to generate a structured feature set.
[0018] As a preferred scheme of the customer data processing method, the modification of the clustering boundary of the state node and the transition edge weight in the state transition graph comprises: marking the high jump gradient region and the multi-region synchronous response region in the structured feature set as key abnormal response points, and writing back to the corresponding state node of each regional state transition graph; and the updated state transition graph forms a dynamic evolution structure based on the feedback of the collaborative state mapping matrix.
[0019] As a preferred scheme of the customer data processing method, the high jump gradient region and the multi-region synchronous response region comprise: the high jump gradient region comprises: in the structured feature set, for each regional linkage state segment, the corresponding task key frame state sequence and the jump amplitude value are extracted; wherein the jump amplitude value is the absolute value of the difference between the adjacent transition probabilities of the state node; in the continuous state sequence, a sliding window mechanism is used to traverse each state sequence of a fixed length, and when the average jump amplitude value in the window exceeds the preset jump threshold value, and the local maximum jump amplitude value is greater than the jump peak value determination threshold, the corresponding state segment is marked as a high jump gradient region; and the multi-region synchronous response region comprises: in the structured feature set, the linkage state segments from different regions are aligned and analyzed, and according to the period normalized time index of the state node, the state node groups in all regions within the same time window are jointly compared; if within the time window, the state nodes of multiple regions are consistent in the jump amplitude value direction, and the jump amplitude values all exceed the cross-region jump consistency threshold, then the state segment corresponding to the time window is marked as a multi-region synchronous response region.
[0020] In a second aspect, the application provides a customer data processing system, comprising: an event cleaning and mapping module for initially cleaning the power consumption event sequence of each regional power grid customer and constructing a state transition graph; a key frame extraction module for constructing a bidirectional state structure based on the state transition graph, extracting an atypical state sequence in each region based on the maximum transition probability path, and labeling it as a regional task key frame state sequence; a collaborative mapping identification module for constructing a collaborative state mapping matrix based on the extracted task key frame state sequence in each region, quantifying the abnormal transition correlation between different regions through a behavior difference vector, and identifying a distributed state abnormal segment with correlation aggregation characteristics; an evolution structure generation module for time sequence encoding the task key frame state sequence corresponding to the state abnormal segment with correlation aggregation characteristics in the collaborative state mapping matrix, and generating a structured feature set representing the customer's phased response behavior; and modifying the clustering boundary of the state node and the transition edge weight in the state transition graph based on the structured feature set, and forming a dynamic evolution expression structure for specific regional behavior abnormalities.
[0021] In a third aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and wherein the computer program instructs the processor to implement the steps of the customer data processing method according to the first aspect of the present application.
[0022] In a fourth aspect, the present application provides a computer readable storage medium storing a computer program, wherein the computer program instructs a processor to implement the steps of the customer data processing method according to the first aspect of the present application.
[0023] The present application has the following beneficial effects: the present application can accurately extract typical behavior stages from a global perspective by constructing a bidirectional state structure oriented to behavior transition rules, and realizes dynamic positioning of key behavior nodes, thereby enhancing the tracing and intervention capabilities for abnormal behaviors; in a multi- regional scenario, a collaborative analysis strategy is used to effectively identify synchronous abnormal patterns across regions, thereby improving the identification accuracy of the behavior commonalities and differences of large-scale power grid user groups; in addition, by uniformly encoding and structuring distributed abnormal features, the expression capability of subsequent data modeling and the real-time performance of system response are greatly enhanced; finally, by feeding back and correcting the original behavior network structure, a dynamic evolution expression mechanism is formed, so that the entire system has the capabilities of continuous learning and self-adaptive optimization, thereby providing strong data support and intelligent protection for the fine customer management of power enterprises. BRIEF DESCRIPTION OF DRAWINGS
[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0025] Figure 1 The flowchart of the customer data processing method.
[0026] Figure 2 The structure diagram of the customer data processing system. DETAILED DESCRIPTION
[0027] In order to make the above-mentioned purposes, features and advantages of the present application more apparent and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0028] In the following description, many specific details are set forth in order to provide a thorough understanding of the present application, but the present application can also be implemented in other ways different from the description, and those skilled in the art can make similar generalizations without departing from the scope of the present application, therefore the present application is not limited to the specific embodiments disclosed below.
[0029] Secondly, the "one embodiment" or "embodiment" referred to herein means a specific feature, structure, or characteristic under consideration or implementation can be included in at least one implementation of the present application. "In one embodiment" appearing in various places in the specification does not all refer to the same embodiment, nor are these embodiments mutually exclusive of other embodiments.
[0030] As described in the above background, the prior art has multiple bottlenecks to be broken through in the aspects of processing regional distributed power customer behavior anomaly identification and structure modeling. On the one hand, typical methods mostly rely on statistical features or frequent behavior segments in a single region, ignoring the time sequence coupling relationship in the evolution process of customer behavior, especially in the identification ability under low-frequency atypical event sequence, which is easy to cause the key behavior segment to be missed. On the other hand, in the multi-regional data fusion and collaborative analysis, the existing method lacks a systematic state structure mapping and cross-regional anomaly association modeling mechanism, which makes it difficult to share and compare behavior characteristics between regions, hindering the identification and expression of large-scale power customer behavior anomaly aggregation effect. In addition, the existing state transition modeling mostly adopts fixed edge weight or rule updating mode, which lacks sensitivity to dynamic changes of events, and it is difficult to effectively reflect the dynamic migration trend of customer response strategy in the time evolution process.
[0031] Therefore, there is an urgent need for a customer data processing method that can face multiple regions of customers, has collaborative perception ability and dynamic evolution modeling ability, to improve the identification accuracy of abnormal behavior segments and the expression depth of behavior response modeling.
[0032] Figure 1 The flowchart of the customer data processing method according to the embodiment of the present application is shown in FIG. 1. As shown in FIG. 1, in the customer data processing method, it includes: Figure 1
[0033] S1: The power consumption event sequence of each regional power grid customer is initially cleaned, and a state transition graph is constructed.
[0034] Firstly, the original power consumption event sequence of each regional power grid customer in a specific period in the past is segmented and sliced according to the timestamp, and a context label index is constructed according to the electricity price period, seasonal load change and important holiday label, which is used to constrain the boundary of the subsequent event aggregation window. For example, independent label regions are constructed in the summer air conditioner high-frequency power consumption period and the Spring Festival low-frequency intermittent power consumption period, which can effectively improve the difference and clarity of state expression, and provide structural guarantee for multi-period power consumption mode identification.
[0035] Further, based on the context label index, the behavior similarity distance of each power consumption event segment is measured, and the local denoising processing is performed in a density-constrained dynamic time warping manner to retain the event mutation points and periodic main feature nodes. Specifically, the process adopts a similarity measurement method based on dynamic time warping (DTW) and density constraint to perform equidistant alignment processing on the time sequence structure difference between event segments, thereby identifying event pairs with highly similar power consumption behaviors. On this basis, smoothing replacement operation is performed on low-frequency fluctuation nodes or isolated power consumption points, while ensuring that the mutation nodes and periodic peak nodes are not weakened.
[0036] Secondly, an embedded spectral clustering-based method is used to compress the state space of the cleaned power consumption event segments, and the event groups are mapped to state nodes; wherein the clustering is associated according to the periodic density peak position and the behavior fluctuation amplitude in the event segment.
[0037] Specifically, in the spectral clustering process, the periodic density peak position (i.e. the maximum point of the concentration of behavior in unit time) and the behavior fluctuation amplitude (i.e. the mean square deviation of continuous event power change) are taken as the dual clustering factors, and two constraint terms of maximum intra-class distance and minimum inter-class interval are set to ensure that the clustering results capture the main behavior distribution characteristics while taking into account the compactness of intra-class behavior and the clarity of inter-class boundary. After clustering, each state node represents a power consumption behavior mode, which can be uniquely represented at the time, frequency and power levels.
[0038] After the state nodes are determined, the state transition probability is calculated according to the time sequence order of adjacent states in the power consumption event segment and the mutation weight, and the state transition graph is constructed.
[0039] Specifically, first, the state nodes in the cleaned event segment are rearranged in the original time stamp order to form a state path chain with time sequence. Then, each pair of adjacent state nodes is traversed in turn, and the transition frequency (i.e. the number of common occurrences between two states) and the mutation intensity (i.e. the difference between the power amplitude of adjacent state nodes is calculated and normalized (such as subtracting the mean and dividing by the standard deviation)) observed in the original event sequence are used to determine the weight of the edge. Specifically, the edge weight is defined as the weighted fusion value of the frequency weight and the mutation amplitude weight, and an adjustable coefficient is set to control the importance of the two types of weights to adapt to different weight requirements of behavior transition frequency and amplitude in different scenarios.
[0040] S2: Construct a bidirectional state structure based on the state transition graph, extract non-typical state sequences in each region based on the maximum transition probability path, and label them as task critical frame state sequences in the region.
[0041] S2.1: For the state node pairs with transition relationship in the state transition graph, the forward and reverse transition probabilities are calculated respectively; wherein the forward transition probability is the state transition probability from the starting state node to the target state node, and the reverse transition probability takes the target state node as the backtracking starting point, extracts the reverse jump path and calculates the corresponding reverse transition frequency, and finally a bidirectional symmetric state transition probability matrix is constructed.
[0042] It should be noted that the traditional method often only focuses on the forward transition relationship, i.e. the transition path from a state to another state, and the basis for construction is the frequency of event sequence change in time sequence. However, in user behavior, some abnormal behaviors are manifested as rarely "backtracking" to the previous state, for example, high-load behavior suddenly appears in off-peak period or power consumption mode similar to weekend appears on weekdays, which is difficult to be captured in forward analysis. Therefore, in the present application, for each pair of state nodes, not only the forward transition probability (i.e. the normalized value of the transition frequency from state A to state B) is recorded, but also the reverse transition probability is calculated, i.e. the possibility of backtracking from state B to state A, to construct a symmetric transition probability structure.
[0043] Specifically, according to the state transition graph constructed in step S1, the frequently occurring state nodes in each region are selected as the backtracking starting point and marked as candidate target state nodes;
[0044] Taking each target state node as the starting point, the reverse traversal along the time sequence in the state transition graph is performed, the backtracking step window and the time span threshold are limited, the possible source of the previous state path segment is extracted, and the reverse transition frequency between the state pairs is recorded; the maximum backtracking step window and the allowed time span threshold are strictly controlled in this search process to avoid introducing invalid paths caused by long-term evolution;
[0045] The frequency of all state pairs in the extracted reverse state path (i.e. the path segment of a state transition to the target state) is counted, and the reverse transition probability value of each state pair is calculated by normalizing the mutation weight index in the path;
[0046] The obtained state transition probability and the calculated reverse transition probability are combined by weighting according to the one-to-one correspondence of the node pairs, to construct a state transition probability matrix with bidirectional symmetry.
[0047] S2.2: For each regional state transition graph, the path cumulative probability is constructed based on the state transition probability product, from each state node to the corresponding periodic terminal state node, and the probability maximum path extraction strategy is used to filter out the atypical path sequence with the maximum cumulative transition probability value and the frequency lower than the historical average value.
[0048] In a specific operation, firstly, a node-by-node traversal is performed on a state transition graph in each region to identify each possible starting state (i.e., a cycle starting state node), and a cycle terminal state is set as a target end point. In the path construction process, instead of only relying on a shortest path or a least jump strategy, a cumulative transition probability of each possible path is taken as a weight basis, and a probability product model is used for dynamic programming type search.
[0049] The product path probability can truly reflect the overall transition credibility of the path, and effectively avoids misjudgment caused by individual high-weight transition nodes. Meanwhile, in the path screening standard, the application further compares the occurrence frequency of the current path in the customer historical behavior database with the average path frequency in the region according to a historical frequency comparison mechanism, eliminates those candidate items which have a high probability but are typical behavior paths, and thus retains the non-typical state transition sequences which have a high transition probability but a low occurrence frequency in a true sense.
[0050] S2.3: Calculate the state jump gradient and the before-and-after path symmetric difference value for each screened non-typical path sequence, and mark the state nodes with continuous high mutation gradient and significant path asymmetry as task critical frames.
[0051] In order to further identify the real key customer behavior change nodes from the non-typical paths, the application takes the state jump gradient and the before-and-after path symmetric difference value as the criterion for task critical frame screening.
[0052] Among them, the state jump gradient is used to measure the transition intensity change rate between continuous state nodes in the non-typical path, that is, to perform a differential amplification process on the mutation degree between states, so as to judge whether a short-time drastic behavior switching occurs in a section; and the before-and-after path symmetric difference value is to compare whether the transition probabilities of the path under the forward and reverse paths remain consistent, if the difference is significant, it indicates that the behavior only exists significantly in a certain direction, has obvious aperiodic characteristics, and is extremely likely to be an event-driven or abnormal induced behavior. In actual implementation, the two are combined in a weighted linear superposition manner to form a joint score, each screened non-typical path sequence is traversed, and the state nodes with a score value higher than a set score threshold are marked as task critical frames, which means that the user may enter an abnormal power consumption phase or a behavior mutation occurs at the state.
[0053] In summary, by constructing a bidirectional state transition structure, the application not only overcomes the technical defects of the existing one-way transition model which cannot capture aperiodic abnormal behaviors, but also constructs a state analysis framework with better interpretability and behavior prediction ability by combining path probability modeling and statistical anomaly detection.
[0054] S3: Construct a collaborative state mapping matrix based on the extracted task-critical frame state sequences in each region, quantify the abnormal transition correlation between different regions through the behavior difference vector, and identify the distributed state abnormal segment with correlation aggregation characteristics.
[0055] S3.1: For the task-critical frame state sequence in each region, construct a behavior feature vector sequence based on the state jump amplitude and time position difference, wherein the jump amplitude is calculated according to the difference in transition probability of the path before and after the key frame, and the time position is represented by a period normalized index.
[0056] Specifically, the state jump amplitude is embodied by measuring the transition probability change of the path before and after the key frame, that is, calculating the difference between the transition probability from the state before the jump to the current key frame state and the transition probability from the key frame state to the next state.
[0057] The time position difference is expressed by a period normalized index, that is, the absolute position of the state node in the period is mapped to the range [0, 1].
[0058] S3.2: Use a vector mapping strategy based on a dynamic time alignment mechanism to perform bidirectional mapping comparison of the behavior feature vector sequences of different regions, construct a behavior similarity matrix, and retain high-similarity state pairs with a similarity higher than a similarity threshold.
[0059] For example, in order to identify similar transition patterns of key frame state sequences between different regions, a vector comparison mechanism compatible with variable length sequences and non-equal length time positions needs to be constructed. For this purpose, the present application uses dynamic time warping (DTW) to perform bidirectional mapping comparison of the behavior feature vector sequences, calculates a local distance measure, and improves the robustness and correlation accuracy of the comparison.
[0060] Based on the above local distance measure, bidirectional DTW mapping is performed to construct a similarity matrix, wherein each element represents the minimum alignment distance between the ith key frame in region A and the jth key frame in region B.
[0061] Further, the alignment distance value is normalized and converted into a similarity score, and a similarity threshold is set. Only high-similarity state pairs with a similarity score greater than or equal to the similarity threshold are retained, which are marked as candidate collaborative state pairs.
[0062] S3.3: Based on the behavior similarity matrix, count the number of times that the high-similarity state pairs appear together in all region combinations; when a high-similarity state pair appears together in a region combination more than a support threshold, it is determined as a cross-regional collaborative state pair, and a collaborative state mapping matrix is constructed, which marks the region combination index and behavior difference indicator vector.
[0063] Each element of the collaborative state mapping matrix contains the index of the collaborative state pair, the regional combination set where the collaborative state pair appears, the average behavior difference index vector of the collaborative state pair, etc.
[0064] S3.4: According to the behavior difference index vector in the collaborative state mapping matrix, the state segment showing similar transition structure in multiple regions is divided into a distributed state abnormal segment with correlation aggregation characteristics, including the following operation steps:
[0065] Filtering continuous high-similarity state pairs from the collaborative state mapping matrix to combine into multiple state segments;
[0066] Comparative analysis of the behavior difference index involved in each state segment to determine whether the performance in different regions is similar; specifically, the average behavior difference vector set involved in each state segment is extracted, the variance is calculated, and compared with the preset behavior difference tolerance threshold. If the variance is less than or equal to the preset behavior difference tolerance threshold, it is determined that the behavior similarity condition is met;
[0067] When the state segment shows similar transition structure in multiple regions and the behavior difference index meets the set threshold, the corresponding state segment is determined as a distributed state abnormal segment with correlation aggregation characteristics.
[0068] S4: Time series encoding of the task critical frame state sequence corresponding to the state abnormal segment with correlation aggregation characteristics in the collaborative state mapping matrix to generate a structured feature set representing the customer's phased response behavior; based on the structured feature set, the clustering boundary and transition edge weight of the state node in the state transition graph are modified to form a dynamic evolution expression structure for specific regional behavior abnormalities.
[0069] It should be noted that after the clustering and positioning of the abnormal segments in the collaborative state mapping matrix are completed, the task critical frame state sequence associated with these abnormal segments needs to be further extracted and converted into a structured feature set in a time series encoding manner, so as to reflect the phased response behavior characteristics of the customer in a specific time period. On this basis, the structured feature set is written back to the original state transition graph to modify the state node clustering boundary and transition edge weight, so that the graph can more effectively express the dynamic evolution characteristics of regional behavior abnormalities.
[0070] S4.1: Based on the distributed state abnormal segment, a set of multi-regional task critical frame state sequences corresponding to the collaborative state mapping matrix is extracted to construct a cross-regional time series transition window and uniformly encode into a multi-regional linkage state segment.
[0071] In a specific operation, first, based on the aforementioned clustering, abnormal fragments with similar state transition patterns and close occurrence times in data distributed in different regions are merged into cross-regional state fragment clusters; then, by setting a unified cross-regional time sequence transition window, key frame state sequences with consistent response structures in each region are time-synchronized to form a set of linked state segments on a unified time axis.
[0072] S4.2: For multi-regional linked state segments, a multi-dimensional time sequence encoding method is used to fuse the transition probability, jump gradient, and regional difference characteristics of the state nodes to generate a structured feature set for revealing the phased characteristics of multi-regional customer behavior responses.
[0073] Specifically, first, for each region, the transition probability between adjacent states is calculated, and a jump gradient sequence is constructed accordingly, with the jump gradient represented by the absolute value of the difference between the transition probabilities; second, to improve cross-regional comparability, the jump characteristics of each region need to be normalized; third, by periodic time indexing, the state nodes are classified and merged according to the task period; finally, the three-tuple features of the state nodes (transition probability, jump gradient, and normalized time index) are combined with the regional label to form a multi-region-multi-dimensional joint structured feature matrix.
[0074] S4.3: Mark the high jump gradient area and multi-regional synchronous response area in the structured feature set as key abnormal response points and write them back to the corresponding state nodes of each regional state transition graph to realize dynamic evolution expression of multi-regional linked abnormal behavior by adjusting the node clustering boundary and transition edge weight.
[0075] Among them, the high jump gradient area includes: in the structured feature set, for each region's linked state segment, the corresponding task key frame state sequence and jump amplitude value are extracted; the jump amplitude value is the absolute value of the difference between the adjacent transition probabilities of the state nodes; in the continuous state sequence, a sliding window mechanism is used to traverse each fixed-length state sequence, when the average jump amplitude value in the window exceeds the preset jump threshold, and the local maximum jump amplitude value is greater than the jump peak value determination threshold, the corresponding state segment is marked as a high jump gradient area. The above operation is used to identify state sequences that have mutated behavior within a certain time period to reflect the dramatic characteristics of local behavior response.
[0076] The multi-region synchronous response region comprises: in the structured feature set, the linkage state segments from different regions are aligned and analyzed, and according to the period normalized time index of the state node, the state node groups in all regions in the same time window are jointly compared; if in the time window, the state nodes of multiple regions are consistent in the direction of the jump amplitude value (that is, the transition probability change sign is the same), and the jump amplitude value exceeds the cross-region jump consistency threshold, then the state segment corresponding to the time window is marked as a multi-region synchronous response region. The marking process is used to determine that multiple regions have consistent behavior patterns at a specific period position, representing the spatio-temporal synchronization feature of the linkage customer behavior response.
[0077] Finally, the above two abnormal regions are written back to the state transition graph of each region as key abnormal response points, and the following dynamic adjustment is performed on the related state nodes in the state transition graph of each region:
[0078] On the one hand, for the mutation state segment appearing in the high jump gradient region, the cluster center distance distribution of the key frame state node involved in the initial clustering is traced back, if the Euclidean distance of the node to the class center is close to the boundary of the maximum intra-class distance constraint value, then the maximum intra-class distance threshold of the current class is dynamically contracted, the class behavior fluctuation amplitude tolerance is compressed, and the discrimination resolution of the nodes around the mutation point is improved; at the same time, if the period density peak position difference between adjacent class centers is less than the preset minimum inter-class interval, then the interval expansion is performed on the clustering boundary between the class centers, so that the behavior clustering trend of the abnormal state segment in the time dimension can be independently represented in the clustering structure, and the identification sensitivity of the state node boundary is enhanced.
[0079] On the other hand, for the state node pair in the high jump gradient region, the transition probability presents a significant transition (that is, the jump amplitude value exceeds the jump threshold) in the local time period, at this time, the edge weight is increased to highlight the decision weight of the transition path under abnormal response; and for the state node group in the multi-region synchronous response region which is in the same period normalized time window and has consistent jump direction, the weight of the corresponding transition edge in each region graph is uniformly improved to construct an edge weight structure with multi-region response synchronization, and the adaptability and expression ability of the model to the spatio-temporal consistent behavior pattern are enhanced.
[0080] After the double feedback adjustment of the clustering boundary and the edge weight, the updated state transition graph of each region will present a dynamic evolution structure: on the one hand, the clustering attribution of the state node is more compact and identifiable around the abnormal response point; on the other hand, the edge weight of the transition path is more sensitive to behavior mutation and has cross-region response coordination, thereby realizing dynamic modeling and evolution expression for multi-region linkage abnormal behavior.
[0081] S4.4: The updated state transition diagram forms a dynamic evolution structure based on collaborative state mapping matrix feedback, and the geographical index of the abnormal response point and the source of the task critical frame are synchronously recorded for subsequent customer response analysis and optimization based on multi-geographical behavior characteristics.
[0082] Further, as shown in Figure 2 the embodiment also provides a customer data processing system, comprising,
[0083] an event cleaning mapping module, configured to perform initial cleaning on the power consumption event sequence of each regional power grid customer, and construct a state transition diagram;
[0084] a key frame extraction module, configured to construct a bidirectional state structure for the state transition diagram, extract an atypical state sequence based on a maximum transition probability path in each region, and label the atypical state sequence as a regional task critical frame state sequence;
[0085] a collaborative mapping identification module, configured to construct a collaborative state mapping matrix based on the extracted task critical frame state sequence in each region, quantize the abnormal transition correlation between different regions through a behavior difference vector, and identify a distributed state abnormal segment with a correlation aggregation characteristic;
[0086] an evolution structure generation module, configured to perform time series coding on the task critical frame state sequence corresponding to the state abnormal segment with the correlation aggregation characteristic in the collaborative state mapping matrix, generate a structured feature set representing the customer's phased response behavior, and correct the clustering boundary and transition edge weight of the state node in the state transition diagram based on the structured feature set, to form a dynamic evolution expression structure for a specific regional behavior anomaly.
[0087] The embodiment also provides a computer device suitable for the customer data processing method, comprising a memory and a processor; the memory is configured to store computer executable instructions, and the processor is configured to execute the computer executable instructions to implement the customer data processing method proposed in the above embodiment.
[0088] The computer device can be a terminal, and the computer device includes a processor, a memory, a communication interface, a display screen and an input device connected by a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with external terminals in a wired or wireless manner. The wireless manner can be achieved by WIFI, an operator network, NFC (near field communication) or other technologies. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen. The input device of the computer device can be a touch layer overlaid on the display screen, or a key, trackball or touchpad arranged on the shell of the computer device. In addition, the input device can be an external keyboard, touchpad or mouse, etc.
[0089] The embodiment also provides a storage medium having a computer program stored thereon, and the program is executed by a processor to implement the client data processing method proposed in the above embodiment.
[0090] To sum up, by constructing a bidirectional state structure oriented to the behavior transition law, the embodiment can accurately extract a typical behavior stage from a global perspective, realizes dynamic positioning of a key behavior node, and thus enhances the tracing and intervention capabilities for abnormal behaviors. In a multi-geographical scenario, a collaborative analysis strategy is used to effectively identify a cross-geographical synchronous abnormal mode, and the identification accuracy for the behavior commonality and difference of a large-scale power grid user group is improved. In addition, by uniformly encoding and structuring distributed abnormal features, the expression capability of subsequent data modeling and the real-time performance of system response are greatly enhanced. Finally, by feeding back and correcting an original behavior network structure, a dynamic evolution expression mechanism is formed, so that the entire system has a continuous learning and self-adaptive optimization capability, and provides strong data support and intelligent protection for fine customer management of power enterprises.
[0091] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and not to limit the present application. Although the present application has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solutions of the present application can be modified or replaced equivalently without departing from the spirit and scope of the technical solutions of the present application, and all should be covered in the scope of the claims of the present application.
Claims
1. A method of processing customer data, characterized by: The method comprises the following steps: initially cleaning the power consumption event sequence of each regional power grid customer, and constructing a state transition graph; constructing a bidirectional state structure for the state transition graph, extracting an atypical state sequence based on the maximum transition probability path in each region, and labeling it as a regional task key frame state sequence; the construction of the bidirectional state structure comprises: for the state node pairs with transition relationship in the state transition graph, the forward and reverse transition probabilities are calculated respectively; wherein the forward transition probability is the state transition probability from the starting state node to the target state node; the reverse transition probability takes the target state node as the backtracking starting point, extracts the reverse jump path and calculates the corresponding reverse transition frequency, and finally constructs the bidirectional symmetric state transition probability matrix; the extraction of the atypical state sequence comprises: for each regional state transition graph, the path cumulative probability is constructed based on the state transition probability product, from each state node to the corresponding period terminal state node, and the probability maximum path extraction strategy is used to filter out the atypical path sequence with the maximum cumulative transition probability value and the frequency lower than the historical average value; constructing a collaborative state mapping matrix based on the task key frame state sequences extracted in each region, quantifying the abnormal transition correlation between different regions through a behavior difference vector, and identifying a distributed state abnormal segment with correlation aggregation characteristics; the construction of the collaborative state mapping matrix comprises: for each regional task key frame state sequence, a behavior feature vector sequence based on state jump amplitude and time position difference is constructed; wherein the jump amplitude is calculated according to the difference between the path transition weights before and after the key frame; conducting bidirectional mapping comparison on the behavior feature vector sequences of different regions, constructing a behavior similarity matrix, and retaining high-similarity state pairs with a similarity higher than a similarity threshold; based on the behavior similarity matrix, the number of times that the high-similarity state pairs commonly appear in all regional combinations is counted; when the common appearance frequency of a high-similarity state pair in a regional combination exceeds a support threshold, it is determined as a cross-regional collaborative state pair, and a collaborative state mapping matrix is constructed, and the regional combination index and behavior difference index vector belonging to it are labeled; the identification of the distributed state abnormal segment with correlation aggregation characteristics comprises: according to the behavior difference index vector in the collaborative state mapping matrix, the state segment with similar transition structure in multiple regions is divided as a distributed state abnormal segment with correlation aggregation characteristics; the task key frame state sequence corresponding to the state abnormal segment with correlation aggregation characteristics in the collaborative state mapping matrix is time series encoded to generate a structured feature set representing the customer's phased response behavior; based on the structured feature set, the clustering boundary and transition edge weight of the state node in the state transition graph are modified to form a dynamic evolution expression structure for specific regional behavior anomaly, and the modification of the clustering boundary and transition edge weight of the state node in the state transition graph comprises: marking the high jump gradient area and multi-regional synchronous response area in the structured feature set as key abnormal response points, and writing back to the corresponding state nodes of each regional state transition graph; The updated state transition diagram forms a dynamic evolution structure based on collaborative state mapping matrix feedback.
2. The client data processing method of claim 1, wherein: The state nodes in the state transition diagram are determined based on customer historical power consumption mode clustering, and the transition edge weights are calculated according to power consumption frequency and behavior mutation degree.
3. The client data processing method of claim 2, wherein: The construction of the state transition diagram includes: An embedded spectral clustering-based method is used to compress the state space of the cleaned power consumption event segments, and the event groups are mapped to state nodes. After the state nodes are determined, the state transition probabilities are calculated according to the time sequence order and mutation weight of adjacent states in the power consumption event segments, and the state transition diagram is constructed.
4. The client data processing method of claim 1, wherein: The calculation of the reverse transition frequency is: Starting from each target state node, traversing the state transition diagram in reverse time order, limiting the backtracking step window and time span threshold, extracting the possible source sequence state path segment, and recording the reverse transition frequency between state pairs; The reverse state path is extracted, and the frequency of all state pairs in the reverse state path is counted, and the reverse transition probability value of each state pair is calculated by normalization.
5. The client data processing method of claim 1, wherein: The division of the state segment presenting similar transition structure in multiple regions includes: Filtering continuous high-similarity state pairs from the collaborative state mapping matrix, and combining them into multiple state segments; Comparative analysis of the behavior difference indicators involved in each state segment is performed to determine whether the performance in different regions is similar; When the state segment shows similar transition structure in multiple regions and the behavior difference indicators meet the set threshold, the corresponding state segment is determined as a distributed state abnormal segment with correlation and aggregation characteristics.
6. The client data processing method of claim 1, wherein: The time sequence coding includes: Based on the distributed state abnormal segment, a set of multi-region task key frame state sequences is extracted from the collaborative state mapping matrix, a cross-region time sequence transition window is constructed, and a multi-region linkage state segment is uniformly coded; For the multi-region linkage state segment, a multi-dimensional time sequence coding method is used to fuse the transition probability, jump gradient and regional difference characteristics of the state nodes to generate a structured feature set.
7. The client data processing method of claim 6, wherein: The high jump gradient area and the multi-region synchronous response area include: The high jump gradient area includes: in the structured feature set, for each regional linkage state segment, the corresponding task key frame state sequence and jump amplitude value are extracted; the jump amplitude value is the absolute value of the difference between the adjacent transition probabilities of the state nodes; in the continuous state sequence, a sliding window mechanism is used to traverse each fixed length state sequence, and when the average jump amplitude value in the window exceeds the preset jump threshold and the local maximum jump amplitude value is greater than the jump peak value determination threshold, the corresponding state segment is marked as a high jump gradient area; The multi-region synchronous response area includes: in the structured feature set, the linkage state segments from different regions are aligned and analyzed, and according to the periodic normalized time index of the state nodes, the state node groups in the same time window in all regions are jointly compared; if in the time window, the state nodes of multiple regions are consistent in the jump amplitude value direction, and the jump amplitude values all exceed the cross-region jump consistency threshold, the state segment corresponding to the time window is marked as a multi-region synchronous response area.
8. A customer data processing system based on the customer data processing method according to any one of claims 1 to 7, characterized by: It also includes: An event cleaning and mapping module is configured to perform initial cleaning on power consumption event sequences of customers in each regional power grid and construct a state transition graph. A key frame extraction module is configured to construct a bidirectional state structure based on the state transition graph, extract an atypical state sequence based on a maximum transition probability path in each region, and label the atypical state sequence as a task key frame state sequence in the region. The construction of the bidirectional state structure includes: for a state node pair having a transition relationship in the state transition graph, a forward transition probability and a reverse transition probability are calculated respectively; the forward transition probability is a state transition probability from a starting state node to a target state node; the reverse transition probability takes the target state node as a backtracking starting point, extracts a reverse jump path, and calculates a corresponding reverse transition frequency, and finally a bidirectional symmetric state transition probability matrix is constructed; the extraction of the atypical state sequence includes: for each regional state transition graph, a path cumulative probability is constructed based on a state transition probability product, a non-typical path sequence having a maximum cumulative transition probability value and a frequency lower than a historical average value is filtered out from each state node to a corresponding periodic terminal state node by using a probability maximum path extraction strategy. A collaborative mapping identification module is configured to construct a collaborative state mapping matrix based on the task key frame state sequences extracted in each region, quantify abnormal transition correlation between different regions by a behavior difference vector, and identify a distributed state abnormal segment having a correlation aggregation characteristic; the construction of the collaborative state mapping matrix includes: For the task key frame state sequence in each region, a behavior feature vector sequence based on state jump amplitude and time position difference is constructed; the jump amplitude is calculated according to a difference degree of path transition weight before and after the key frame; The behavior feature vector sequences of different regions are compared by bidirectional mapping, a behavior similarity matrix is constructed, and a high-similarity state pair having a similarity higher than a similarity threshold is retained; Based on the behavior similarity matrix, the number of times that the high-similarity state pair appears together in all regional combinations is counted; when the number of times that the high-similarity state pair appears together in the regional combinations exceeds a support threshold, the high-similarity state pair is determined as a cross-regional collaborative state pair, and a collaborative state mapping matrix is constructed, and a regional combination index and a behavior difference index vector belonging to the high-similarity state pair are labeled; The identification of the distributed state abnormal segment having the correlation aggregation characteristic includes: According to the behavior difference index vector in the collaborative state mapping matrix, a state segment having a similar transition structure in multiple regions is divided as the distributed state abnormal segment having the correlation aggregation characteristic; An evolution structure generation module is configured to perform time sequence coding on a task key frame state sequence corresponding to a state abnormal segment having a correlation aggregation characteristic in the collaborative state mapping matrix, generate a structured feature set representing a customer's phased response behavior, and correct a clustering boundary and a transition edge weight of a state node in the state transition graph based on the structured feature set, to form a dynamic evolution expression structure facing a specific regional behavior anomaly; the correction of the clustering boundary and the transition edge weight of the state node in the state transition graph includes: Mark the high jump gradient region and the multi-region synchronous response region in the structured feature set as key abnormal response points, and write back to the corresponding state nodes of the state transition graph of each region; The updated state transition graph forms a dynamic evolution structure based on the feedback of the collaborative state mapping matrix.
Citation Information
Patent Citations
Association processing method and device of power grid data, computer equipment and storage medium
CN116976507A
Power grid abnormal user transformer processing method and system
CN117541224A