Power supply service risk identification system based on big data analysis

By constructing a power supply service risk identification system based on big data analysis, the problems of single data and incomplete assessment in power supply service risk identification have been solved. It has realized the comprehensive analysis of multi-source data and intelligent risk identification, thereby improving the level of intelligence in power supply service risk management.

CN119809319BActive Publication Date: 2026-08-25STATE GRID HUBEI MARKETING SERVICE CENT (MEASUREMENT CENT)
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411839198.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2026-08-25
Estimated Expiration
2044-12-13

AI Technical Summary

Technical Problem

Existing power supply service risk identification technologies suffer from problems such as limited data sources, simplistic data processing methods, incomplete risk assessment indicators, and lagging early warning mechanisms, making it difficult to comprehensively and accurately identify potential risks in power supply services.

Method used

A power supply service risk identification system based on big data analysis is constructed. By collecting multi-source data (work order data, feedback data, and public opinion data), the system performs data preprocessing, fusion, and standardization. Features are extracted using LSTM networks, multi-level attention mechanisms, and graph convolutional networks. A two-layer risk assessment index system and dynamic threshold model are established to achieve intelligent risk identification and early warning.

Benefits of technology

It enables multi-dimensional quantitative assessment and accurate identification of power supply service risks, timely detection of potential risks and provision of handling suggestions, thereby improving the level of intelligence in power supply service risk management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119809319B_ABST
    Figure CN119809319B_ABST
Patent Text Reader

Abstract

The application provides a power supply service risk identification system based on big data analysis, relates to the technical field of power supply services, and comprises the following: a data acquisition module, which is used for acquiring multi-source data related to power supply services, wherein the multi-source data comprises work order data, feedback data and public opinion data; a data processing module, which is used for pre-processing the multi-source data to obtain standardized data; a data fusion module, which is used for fusing the standardized data to obtain fused data; a risk identification module, which is used for analyzing the fused data, identifying the risk types in the power supply services by establishing a power supply service risk evaluation index set, and determining the risk grades; and an early warning output module, which is used for generating risk early warning information according to the risk identification results, wherein the risk early warning information comprises the risk types, the risk grades, the influence ranges and disposal suggestions. The application can comprehensively identify, accurately evaluate and timely warn the power supply service risks, thereby effectively improving the power supply service risk management level.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power supply service technology, and in particular to a power supply service risk identification system based on big data analysis. Background Technology

[0002] With the rapid development of the economy and society, electricity supply has become a fundamental guarantee for the operation of modern society. Electricity supply service, as an important function of power companies, is directly related to the daily lives of millions of households and the normal operation of enterprises. Electricity supply service involves multiple stages, including electricity application, fault repair, and complaint handling. The service process is complex and influenced by numerous factors, placing high demands on service quality management.

[0003] Power supply service risk identification is a crucial means of ensuring service quality. Currently, power supply companies primarily identify service risks by collecting work order data and customer feedback, employing manual analysis and simple statistical methods. This risk identification mechanism is of great significance for preventing service quality issues and improving service levels.

[0004] However, existing power supply service risk identification technologies have several shortcomings: First, the data sources are limited, making it difficult to comprehensively reflect service quality; second, the data processing methods are simplistic, failing to fully leverage the value of the data; third, the risk assessment indicator system is incomplete, lacking a multi-dimensional evaluation of service quality; and finally, the risk early warning mechanism is relatively lagging, making it difficult to promptly identify and address potential risks. These problems severely restrict the effectiveness of power supply service risk management, necessitating the exploration of new risk identification methods. Summary of the Invention

[0005] In view of this, the present invention proposes a power supply service risk identification system based on big data analysis. By collecting and analyzing multi-source data such as work order data, feedback data, and public opinion data, the system can achieve comprehensive identification, accurate assessment, and timely early warning of power supply service risks, thereby effectively improving the level of power supply service risk management and providing decision support for improving the quality of power supply services.

[0006] The technical solution of this invention is implemented as follows: This invention provides a power supply service risk identification system based on big data analysis, comprising:

[0007] The data acquisition module is used to collect multi-source data related to power supply services, including work order data, feedback data, and public opinion data.

[0008] The data processing module is used to preprocess multi-source data to obtain standardized data;

[0009] The data fusion module is used to fuse standardized data to obtain fused data;

[0010] The risk identification module is used to analyze the fused data, identify the types of risks in the power supply service, and determine the risk level by establishing a set of power supply service risk assessment indicators.

[0011] The early warning output module is used to generate risk warning information based on the risk identification results. The risk warning information includes: risk type, risk level, scope of impact, and handling recommendations.

[0012] Based on the above technical solution, preferably, the data processing module includes:

[0013] The data cleaning unit is used to detect and process outliers in multi-source data. It uses an outlier detection algorithm based on local outlier factors to identify outlier data points and corrects outliers using a sliding window midpoint filling method.

[0014] The data standardization unit is used to standardize the cleaned data. It adopts an adaptive standardization method based on the data distribution characteristics to transform data with different dimensions into a unified numerical range, thus obtaining standardized data.

[0015] Based on the above technical solution, the preferred adaptive standardization method in the data standardization unit is as follows:

[0016] For a data sequence Y, calculate its distribution characteristic parameters α and β:

[0017]

[0018] β = median(I) + σ × skew(I)

[0019] In the formula, Q3 is the upper quartile, Q1 is the lower quartile, m is the sample size; median(Y) is the median, σ is the standard deviation, and skew(Y) is the skewness.

[0020] Standardize the data sequence Y based on the distribution characteristic parameters α and β:

[0021]

[0022] In the formula, ε is a smoothing factor used to prevent the denominator from being zero.

[0023] Based on the above technical solution, preferably, in the data fusion module, a data registration algorithm is used to perform spatiotemporal alignment and feature fusion of multi-source data to achieve spatiotemporal consistency alignment of multi-source data. The calculation formula of the data registration algorithm is as follows:

[0024]

[0025]

[0026] In the formula, A(p,t) represents the alignment result of position p at time t; D i (p,t) represents the observation of the i-th data source at location p and time t; d(p,p i () represents the reference position p between position p and data source i. i Spatial distance between; |tt i | represents the time t and the sampling time t of data source i. i The time difference between them; ω i Q represents the weight coefficient of the i-th data source; i Δs represents the spatiotemporal resolution coefficient of data source i. i Indicates the spatial resolution of data source i; Δt i This represents the temporal resolution of data source i; a1, a2, and γ are adjustment parameters used to control spatial decay, temporal decay, and resolution effects.

[0027] Based on the above technical solution, preferably, the risk identification module includes:

[0028] The feature extraction unit is used to extract service quality features from the fused data;

[0029] The risk assessment unit is used to construct a risk assessment indicator system and calculate the risk assessment indicator values ​​for power supply services based on service quality characteristics.

[0030] The risk classification unit is used to establish a dynamic threshold model based on historical data statistics, determine the risk level according to the power supply service risk assessment index value and the dynamic threshold model, determine the risk type based on the feature combination pattern, and output the risk identification result.

[0031] Based on the above technical solution, the preferred execution process of the feature extraction unit is as follows:

[0032] Extracting time-series features from work order data:

[0033] The work order data is segmented using a sliding time window W(t,Δt), with the window length Δt adaptively adjusted according to the business characteristics: Δt = max(μ d ·(1+σ d ),Δt min ), where μ d σ represents the average processing time for historical work orders. d Δt represents the standard deviation. min This is the lower limit of the window length;

[0034] Construct an LSTM network containing attention-enhanced memory units, and control the information flow through forget gates, input gates, and output gates to extract temporal features f. τ =[fτ1 ,f τ2 ,f τ3 ], where f τ1 As an efficiency trend characteristic, through f τ2 For the characteristics of business volume fluctuation, f τ3 Features of time series patterns;

[0035] Semantic feature extraction is performed on the feedback data:

[0036] Construct a multi-level attention mechanism that includes word-level, sentence-level, and document-level attention:

[0037] α w =softmax(v w ·tanh(W w ·h w ))

[0038] α s =softmax(v s ·tanh(W s ·h s ))

[0039] α d =softmax(v d ·tanh(W d ·h d ))

[0040] In the formula, α w The word-level attention weights have a value range of [0,1], and their sum is 1 within each sentence; α s α represents the sentence-level attention weights, with a range of [0,1] and an in-document sum of 1; d h represents the document-level attention weights, with a value range of [0,1] and a sum of 1. w This is a word-level hidden layer representation, with the dimension being the word embedding dimension; h s This is a sentence-level hidden layer representation, with the dimension being the sentence encoding dimension; h d This is a document-level implicit representation, with the dimension being the document encoding dimension; W w W s W d Here is the attention weight matrix for each layer; v w v s v d These are the attention vectors for each layer;

[0041] By calculating the attention weights of each layer, key information is identified, and semantic features f are extracted. s =[f s1 ,f s2 ,f s3 ], where f s1 For the theme distribution characteristics, f s2For keyword features, f s3 Emotional characteristics;

[0042] Feature extraction from public opinion data:

[0043] Constructing multimodal features of public opinion texts includes: extracting text features using a multi-level attention mechanism, extracting propagation network features through a graph convolutional network (GCN), and analyzing public opinion evolution using a time decay attention mechanism to extract temporal features;

[0044] Calculate and analyze multimodal features to extract public opinion features f m =[f m1 ,f m2 ,f m3 ], where f m1 As a characteristic of public opinion influence, f m2 As a characteristic of public opinion evolution, f m3 This is a feature for public opinion clustering;

[0045] Interaction feature extraction based on multi-source data:

[0046] Construct a data association graph G(V,E), where the node set V contains work order data, feedback data, and public opinion data, and the edge set E represents the association between entities;

[0047] The association strength between nodes is calculated by weighted summation based on content similarity, temporal relevance, and spatial relevance, and an association strength matrix is ​​constructed.

[0048] Interaction features f are extracted based on the association strength matrix and data association graph. c =[f c1 ,f c2 ,f c3 ], where f c1 For relational features, f c2 For propagation characteristics, f c3 For association features;

[0049] For time series features f τ Semantic features f s Public opinion characteristics f m Interaction features f c The system performs standardization and determines the fusion weights using an attention mechanism to obtain the final service quality feature F.

[0050] Based on the above technical solution, the preferred risk assessment indicator system includes two levels of indicators. The first level of indicators includes service response indicators, service quality indicators, user satisfaction indicators, and social impact indicators. The second level of indicators are the next level of indicators set for the first level of indicators. Service response indicators include average response time and first response timeliness rate. Service quality indicators include problem resolution rate and service standardization. User satisfaction indicators include complaint rate and negative review rate. Social impact indicators include public opinion spread and negative impact degree.

[0051] Based on the above technical solution, the preferred process for calculating the power supply service risk assessment index value based on service quality characteristics is as follows:

[0052] Construct a feature mapping function to calculate the second-layer index value vector I:

[0053] I=σ(W·F+b)·(1+tanh(F T ·M·F))

[0054] Where σ is the sigmoid function; F is the service quality feature; W is the feature mapping matrix; M is the feature interaction matrix; and b is the bias vector.

[0055] Calculate the power supply service risk assessment index value R based on the second-level index value vector I:

[0056] R=ψ(I)·(1+η·‖ΔI‖)

[0057] In the formula, ψ(I) is a nonlinear combination function used to integrate the index values ​​of various dimensions, η is a time-varying parameter, ΔI is the change in the index value, and ‖·‖ represents the norm operation.

[0058] Based on the above technical solution, preferably, the dynamic threshold model in the risk classification unit is established through the following steps:

[0059] A time-series feature sequence is constructed based on historical data, and the time-dependent feature z(t) is extracted using a Long Short-Term Memory (LSTM) network.

[0060] Calculate the adaptive factor λ(t):

[0061]

[0062] In the formula, u1 and u2 are adjustment parameters;

[0063] Calculate the dynamic threshold boundaries T1(t) and T2(t), where T1(t) is the threshold value. <T2(t):

[0064] T1(t)=μ his -λ(t)·σ his

[0065] T2(t) = μ his + λ(t)·σ his

[0066] Where, μ his is the mean value of the historical risk assessment index value, and σ his is the standard deviation of the historical risk assessment index value;

[0067] Determine the risk level according to the power supply service risk assessment index value R:

[0068] When R < T1(t), it is determined as low risk;

[0069] When T1(t) ≤ R < T2(t), it is determined as medium risk;

[0070] When R ≥ T2(t), it is determined as high risk.

[0071] On the basis of the above technical solution, preferably, in the risk classification unit, the feature combination mode represents the feature distribution law corresponding to different risk types, including the combination relationship and time series evolution characteristics of each dimension feature in the service quality feature F;

[0072] For each moment t in the historical data, construct a feature pattern matrix M(t):

[0073] M(t) = [F(t), ΔF(t), δF(t)]

[0074] Where, F(t) is the historical service quality feature at moment t, ΔF(t) is the change amount of the feature, and δF(t) is the change rate of the feature;

[0075] Based on the sequence of feature pattern matrices {M(t)} constructed from historical data, use the density clustering algorithm for analysis, cluster the patterns with similar feature distributions and evolution laws into one category, and each category corresponds to a specific risk type, thus forming a risk feature pattern library;

[0076] Construct a feature pattern matrix M(t0) for the current moment t0, and determine the specific risk type by calculating the similarity between M(t0) and each type of pattern in the risk feature pattern library. The similarity calculation uses a non - linear measurement method considering feature importance.

[0077] The present invention has the following beneficial effects compared with the prior art:

[0078] (1) By constructing a power supply service risk identification system based on big data analysis, the present invention realizes the comprehensive analysis of multi - source data of power supply services and the intelligent identification of risks. Compared with the prior art, it can more comprehensively and accurately discover potential risks in power supply services, and timely give risk warnings and disposal suggestions, effectively improving the intelligent level of power supply service risk management;

[0079] (2) The data processing module of this invention, by introducing an anomaly detection algorithm based on local anomaly factors and a sliding window midpoint filling method, can effectively identify and correct outliers in multi-source data, ensuring the integrity and accuracy of the data. Simultaneously, it employs an adaptive standardization method based on data distribution characteristics to convert data of different dimensions into a unified numerical range, ensuring the standardization and consistency of the data processing results.

[0080] (3) This invention achieves spatiotemporal consistency alignment of multi-source data through a data registration algorithm, taking into account the influence of factors such as spatial distance, time difference and data source resolution, solving the spatiotemporal inconsistency problem in the fusion of multi-source heterogeneous data and improving the accuracy of data fusion;

[0081] (4) This invention extracts the temporal features of work order data through a sliding time window, extracts the semantic features of feedback data through a multi-level attention mechanism, and extracts the multimodal features of public opinion data by combining graph convolutional networks and time decay attention mechanisms. This enables the extraction of key feature information from different data sources. At the same time, the interaction feature extraction method based on data association graphs further explores the association characteristics between multi-source data, and the service quality features generated in the end have higher expressive power and discriminative power.

[0082] (5) This invention adopts a two-layer risk assessment index system and a nonlinear combination assessment method. Combined with feature mapping function and interaction matrix, it realizes multi-dimensional quantitative assessment of power supply service risk. At the same time, the risk classification module is based on dynamic threshold model and feature combination mode. Combined with the time series features of historical data and density clustering algorithm, it can accurately determine the risk type and risk level, thereby realizing accurate identification and classification of power supply service risk. Attached Figure Description

[0083] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0084] Figure 1 This is a system framework diagram of an embodiment of the present invention. Detailed Implementation

[0085] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0086] like Figure 1 As shown, this invention provides a power supply service risk identification system based on big data analysis, including:

[0087] The data acquisition module is used to collect multi-source data related to power supply services, including work order data, feedback data, and public opinion data.

[0088] The data processing module is used to preprocess multi-source data to obtain standardized data;

[0089] The data fusion module is used to fuse standardized data to obtain fused data;

[0090] The risk identification module is used to analyze the fused data, identify the types of risks in the power supply service, and determine the risk level by establishing a set of power supply service risk assessment indicators.

[0091] The early warning output module is used to generate risk warning information based on the risk identification results. The risk warning information includes: risk type, risk level, scope of impact, and handling recommendations.

[0092] Specifically, in one embodiment of the present invention, the data acquisition module is responsible for collecting multi-source data related to power supply services, including work order data, feedback data, and public opinion data. In practice, the data acquisition module interfaces with the power supply service management system to obtain relevant data information in real time.

[0093] The work order data originates from the business processing system within the power supply service management system. It includes various types of work orders such as electricity connection applications, fault repair reports, and business inquiries. Each work order record contains fields such as work order number, user information, work order type, work order content, acceptance time, response time, processing process, completion time, and processing result. The data acquisition module periodically synchronizes the latest work order information by calling the data interface of the business processing system and tracks the work order status in real time.

[0094] Feedback data includes two parts: user evaluation data and complaint data. User evaluation data comes from the service evaluation system and includes information such as service rating, evaluation content, and evaluation time. Complaint data comes from the complaint handling system and includes information such as complaint content, complaint type, complaint time, processing process, and processing result. The data collection module collects user evaluation and complaint information in real time through the data interfaces of the service evaluation system and the complaint handling system.

[0095] Public opinion data is sourced from external channels such as social media platforms, news websites, and government complaint platforms. This includes news reports, social media discussions, and online complaints related to power supply services. Each piece of public opinion data contains fields such as information content, publication time, publication platform, dissemination scope, and interaction data. The data collection module continuously monitors and collects relevant public opinion information through web crawling technology and third-party data interfaces.

[0096] During the data collection process, the data collection module also needs to ensure the timeliness and completeness of the data. For work order data and feedback data, a near real-time collection method is adopted, with a fixed collection cycle set, such as 5 minutes, and data is synchronized regularly. For public opinion data, a real-time monitoring method is adopted, and data is collected immediately once relevant information is detected. At the same time, a preliminary integrity check is performed on the collected data to ensure the existence of necessary fields and the correctness of basic format.

[0097] In addition, the data acquisition module also needs to uniformly identify and timestamp the collected data. Each data entry is assigned a unique identifier, and metadata such as the data acquisition time and data source are recorded.

[0098] Specifically, in one embodiment of the present invention, the data processing module includes:

[0099] The data cleaning unit is used to detect and process outliers in multi-source data. It uses an outlier detection algorithm based on local outlier factors to identify outlier data points and corrects outliers using a sliding window midpoint filling method.

[0100] The data standardization unit is used to standardize the cleaned data. It adopts an adaptive standardization method based on the data distribution characteristics to transform data with different dimensions into a unified numerical range, thus obtaining standardized data.

[0101] In this embodiment, the data cleaning unit uses an anomaly detection algorithm based on local anomaly factors to detect and process outliers in multi-source data.

[0102] In specific implementation, first calculate the local density between data points, and identify abnormal data points by comparing the ratio of the local density of a data point to that of its neighborhood points. For work order data, mainly detect outliers in numerical fields such as response time and processing duration; for feedback data, mainly detect outliers in rating data; for public opinion data, mainly detect outliers in statistical indicators such as propagation volume and interaction volume. After identifying the abnormal data points, use the sliding window median filling method to correct the outliers. The specific calculation process is as follows:

[0103] Calculate the k-distance and k-distance neighborhood:

[0104] k-dist(p) = d(p,o)

[0105] Where p is the point to be detected, o is the k-th nearest neighbor of p, d(p,o) is the Euclidean distance between the two points, and k is the preset neighborhood parameter.

[0106] Calculate the reachability distance:

[0107] reach-dist(p,o) = max{k-dist(o), d(p,o)}

[0108] Where reach-dist(p,o) is the reachability distance of point p relative to point o, k-dist(o) is the k-distance of point o, and d(p,o) is the Euclidean distance between point p and point o.

[0109] Calculate the local reachability density:

[0110] lrd(p) = 1 / (∑reach-dist(p,o) / |N k (p)|)

[0111] Where lrd(p) is the local reachability density of point p, N k (p) is the set of points in the k-distance neighborhood of p, and |N k (p)| is the number of neighborhood points.

[0112] Calculate the local outlier factor:

[0113] LOF(p) = (∑(lrd(o) / lrd(p))) / |N k (p)|

[0114] Where LOF(p) is the local outlier factor of point p, lrd(o) is the local reachability density of neighborhood point o, and lrd(p) is the local reachability density of point p. Abnormality determination threshold: When 1 < LOF(p) < 1.5, it is determined as a weak outlier; when LOF(p) ≥ 1.5, it is determined as a strong outlier; when LOF(p) ≈ 1, it is determined as a normal point.

[0115] Sliding window center fill:

[0116] For time series data Y(t), within the time window [tw, t+w]:

[0117] Y′(t)=median({Y(i)|(tw)≤i≤(t+w)})

[0118] Where w is the half-width of the window, Y(i) represents the data value of the time series at time i, {Y(i)|(tw)≤i≤(t+w)} represents the set of all data points within the time window, Y′(t) is the data value after outlier correction, and median() is the function to calculate the median.

[0119] This method sets a fixed-size time window, calculates the median of the data within the window, and uses it to replace outliers. This method can maintain the temporal characteristics of the data and has strong robustness to outliers.

[0120] In this embodiment, the adaptive standardization method proceeds as follows:

[0121] For a data sequence Y, calculate its distribution characteristic parameters α and β:

[0122]

[0123] β = median(I) + σ × skew(I)

[0124] In the formula, Q3 is the upper quartile, Q1 is the lower quartile, m is the sample size; median(Y) is the median, σ is the standard deviation, and skew(Y) is the skewness.

[0125] Standardize the data sequence Y based on the distribution characteristic parameters α and β:

[0126]

[0127] In the formula, ε is a smoothing factor used to prevent the denominator from being zero.

[0128] This embodiment employs an adaptive standardization method based on data distribution characteristics to transform data of different dimensions into a unified numerical range. This method can automatically adjust the standardization parameters according to the distribution characteristics of the data, adapt to the distribution characteristics of different types of data, and ensure the rationality of the standardization results.

[0129] In the specific implementation process, the data processing module also needs to consider the characteristics of different types of data. For time-related data in work order data, such as response time and processing time, standardization is performed using relative time difference; for rating data in feedback data, standardization is performed using a rating interval mapping method; and for statistical indicators in public opinion data, logarithmic transformation is performed before standardization. This classification and processing approach ensures that different types of data are comparable after standardization.

[0130] Specifically, in one embodiment of the present invention, the data fusion module employs a data registration algorithm to perform spatiotemporal alignment and feature fusion on multi-source data, thereby achieving spatiotemporal consistency alignment of the multi-source data. The specific implementation process of the data registration algorithm is as follows:

[0131] First, the algorithm achieves spatiotemporal alignment of multi-source data by calculating formula A(p,t) based on the alignment results of position and time. The formula is:

[0132]

[0133] In the formula, A(p,t) represents the alignment result of position p at time t; D i (p,t) represents the observation at location p and time t of the i-th data source, which is the basic input of the original data; d(p,p i () represents the reference position p between position p and data source i. i The spatial distance between them is used to measure spatial correlation; |tt i | represents the time t and the sampling time t of data source i. i The time difference between them is used to measure time correlation; ω i This represents the weight coefficient of the i-th data source, used to adjust the importance of different data sources; It is a spatial distance decay term, which decays exponentially as spatial distance increases; It is the time-distance decay term, which decays exponentially as the time difference increases; a1 and a2 are adjustment parameters used to control the degree of spatial decay and time decay.

[0134] This embodiment introduces time and spatial attenuation terms to achieve spatiotemporal alignment of data. For example, a substation failure may affect multiple surrounding residential areas, and complaints may be located in different parts of the affected area. The algorithm uses a spatial attenuation function to reasonably establish the correlation between data from different locations; for For example, if a power outage occurs in a certain area, the work order system will record it immediately. However, related complaints and public opinion may emerge in succession. The algorithm uses a time decay function to reasonably correlate these data that differ in time.

[0135] Secondly, the algorithm introduces the spatiotemporal resolution coefficient Q.i This is to adjust for the impact of different data sources. The formula for calculating the spatiotemporal resolution coefficient is:

[0136]

[0137] In the formula, Δs i Δt represents the spatial resolution of data source i, reflecting the level of detail of the data in the spatial dimension; i γ represents the time resolution of data source i, reflecting the sampling frequency of the data in the time dimension; γ is an adjustment parameter used to control the degree of influence of resolution on the final result.

[0138] The spatiotemporal resolution of the data source referred to in this embodiment is an indicator describing the level of detail of the data source in the time and spatial dimensions, Δt. i Indicates the time interval for data collection or updating; Δs i This indicates the spatial positioning accuracy of the data.

[0139] The data registration algorithm provided in this embodiment performs weighted fusion of multi-source data by considering spatial distance attenuation, temporal distance attenuation, and the spatiotemporal resolution of the data sources. The spatial distance attenuation and temporal distance attenuation terms ensure that data that is closer in distance and has a smaller time difference has greater weight. The spatiotemporal resolution coefficient further adjusts the influence of different data sources, allowing data sources with higher resolution to play a greater role in the fusion process.

[0140] In this way, the algorithm achieves spatiotemporal consistency alignment of multi-source data, unifying work order data, feedback data, and public opinion data under the same spatiotemporal reference system. This alignment method considers not only the spatial location and timestamp of the data but also the characteristics of the data source (such as resolution), thereby ensuring the accuracy and reliability of the fusion result. The fused data retains important information from each data source while achieving a unified data representation.

[0141] Specifically, in one embodiment of the present invention, the risk identification module includes:

[0142] The feature extraction unit is used to extract service quality features from the fused data;

[0143] The risk assessment unit is used to construct a risk assessment indicator system and calculate the risk assessment indicator values ​​for power supply services based on service quality characteristics.

[0144] The risk classification unit is used to establish a dynamic threshold model based on historical data statistics, determine the risk level according to the power supply service risk assessment index value and the dynamic threshold model, determine the risk type based on the feature combination pattern, and output the risk identification result.

[0145] In this embodiment, the feature extraction unit is responsible for extracting service quality features from the fused data. The specific implementation includes four main parts: temporal feature extraction from work order data, semantic feature extraction from feedback data, feature extraction from public opinion data, and interactive feature extraction from multi-source data.

[0146] Extracting time-series features from work order data:

[0147] The work order data is segmented using a sliding time window W(t,Δt), with the window length Δt adaptively adjusted according to the business characteristics: Δt = max(μ d ·(1+σ d ),Δt min ), where μ d σ represents the average processing time for historical work orders. d Δt represents the standard deviation. min This is the lower limit of the window length;

[0148] Construct an LSTM network containing attention-enhanced memory units:

[0149] c t =f t ⊙c t-1 +i t ⊙tanh(W c ·[h t-1 ,x t ]+b c )

[0150] h t =o t ⊙tanh(c t )

[0151] a t =softmax(W a ·tanh(W h ·h t ))

[0152] Among them, c t c represents the current state of the memory unit, with the dimension being the hidden layer size; t-1 The state of the memory unit at the previous moment; h t This is the hidden layer output at the current moment, with the same dimension as c. t h t-1 x is the hidden layer output from the previous time step; t f is the input vector at the current time step; t The output of the forget gate has a value range of [0,1] and controls the retention of historical information; i t The input gate output has a value range of [0,1] and controls the reception of new information; tThis is the output gate, with a value range of [0,1], for outputting control information; W c b is the weight matrix of the memory units; c For memory cell bias vector; a t For attention weights at each time step; W a W is the attention weight matrix. h This is the hidden layer transformation matrix.

[0153] The information flow is controlled by a forget gate, an input gate, and an output gate to extract the temporal features f. τ =[f τ1 ,f τ2 ,f τ3 ].

[0154] Among them, f τ1 To identify efficiency trends, an LSTM network is used to analyze the work order processing time sequence. The calculation method is as follows: First, the work order is segmented according to a time window W(t,Δt), and the average processing time within each window is calculated. Then, the LSTM network is used to extract the trend of processing efficiency changes, and a trend feature vector is output to reflect the overall trend of service efficiency. τ2 To analyze the time distribution characteristics of work order quantity to reflect the fluctuation characteristics of business volume, the following steps are taken: The number of work orders within each time window is counted, the fluctuation amplitude between adjacent windows is calculated, periodic fluctuation patterns are extracted, and a fluctuation feature vector is output to reflect the changing patterns of business volume. τ3 To identify the timing pattern of work order processing, the system extracts the work order status transition sequence, analyzes the time characteristics of the processing flow, identifies the abnormal handling pattern, and outputs the pattern feature vector to reflect the standardization of the service process.

[0155] Semantic feature extraction is performed on the feedback data:

[0156] Construct a multi-level attention mechanism that includes word-level, sentence-level, and document-level attention:

[0157] α w =softmax(v w ·tanh(W w ·h w ))

[0158] α s =softmax(v s ·tanh(W s ·h s ))

[0159] α d =softmax(v d ·tanh(W d ·h d ))

[0160] In the formula, α w The word-level attention weights have a value range of [0,1], and their sum is 1 within each sentence; α s α represents the sentence-level attention weights, with a range of [0,1] and an in-document sum of 1; d h represents the document-level attention weights, with a value range of [0,1] and a sum of 1. w This is a word-level hidden layer representation, with the dimension being the word embedding dimension; h s This is a sentence-level hidden layer representation, with the dimension being the sentence encoding dimension; h d This is a document-level implicit representation, with the dimension being the document encoding dimension; W w W s W d Here is the attention weight matrix for each layer; v w v s v d These are the attention vectors for each layer;

[0161] By calculating the attention weights of each layer, key information is identified, and semantic features f are extracted. s =[f s1 ,f s2 ,f s3 ], where f s1 For the theme distribution characteristics, f s2 For keyword features, f s3 For emotional characteristics; specifically, f s1 Extraction based on multi-level attention mechanism: word-level attention α w Keyword identification, sentence-level attention α s Extracting topic sentences, document-level attention α d Determine topic weights and output topic distribution vectors to reflect the topic composition of the feedback content; f s2 Extracting keywords through a word-level attention mechanism: calculating word importance scores, extracting high-frequency keywords, analyzing word co-occurrence relationships, and outputting keyword feature vectors to reflect the core content of the feedback; s3 Sentiment analysis based on multi-layer attention: word-level sentiment polarity recognition, sentence-level sentiment intensity calculation, document-level sentiment tendency judgment, outputting sentiment feature vectors that reflect the user's emotional attitude.

[0162] Feature extraction from public opinion data:

[0163] Constructing multimodal features for public opinion texts includes: extracting text features using a multi-level attention mechanism, extracting propagation network features through a graph convolutional network (GCN), and analyzing public opinion evolution using a time-decaying attention mechanism to extract temporal features. Specifically, the multi-level attention mechanism reuses the multi-level attention mechanism constructed during semantic feature extraction. When extracting propagation network features, a propagation network graph can be constructed, with users participating in the propagation as nodes, the propagation relationships between users as edges, and the connection relationships between nodes constructing an adjacency matrix to form the propagation network graph. Then, GCN is used to extract propagation network features. The temporal attention calculation is as follows:

[0164] α t =softmax(λ(t)·v) t ·tanh(W t ·z t ))

[0165]

[0166] In the formula, λ(t) is the time decay function, and v t W is the temporal attention vector. t Let z be the time-series weight matrix. t For the time-series hidden state, t0 is the current time, t is the historical time point, T is the time window length, and θ is the decay coefficient.

[0167] Calculate and analyze multimodal features to extract public opinion features f m =[f m1 ,f m2 ,f m3 ], where f m1 =[Ic, Ip, Ic·Ip] represents the public opinion influence feature, where Ic is the content influence, calculated by weighting and summing the text sentiment intensity, normalized text length value, and content importance score; Ip is the dissemination influence, calculated by weighting and summing the dissemination scope, dissemination frequency, and dissemination depth; f m2 = [Ps, Vs, Ts] represents the evolution characteristics of public opinion, Ps = [p1, p2, p3, p4] represents the stage characteristics, indicating the characteristic vectors of public opinion in the four stages of emergence, spread, peak, and decline, Vs = dN / dt·τ(t) represents the velocity characteristics, dN / dt represents the speed of public opinion propagation, τ(t) represents the time decay factor, and Ts = ∑(a t ·his t ) represents the trend characteristic, a t For temporal attention weights, his t f is the historical state vector; m3= [Gs, Cs, Ds] represents the public opinion clustering features. Gs is the structural feature, extracted from the output of the last layer of GCN. Cs is the clustering feature, calculated by using spectral clustering. Ds = |Ec| / (|Vc|·(|Vc|-1)) is the density feature. |Ec| is the number of edges in the subgraph, and |Vc| is the number of nodes in the subgraph. Here, a subgraph refers to a local network structure in the propagation network graph, representing a group of interconnected users and their propagation relationships under a certain public opinion theme or event.

[0168] Interaction feature extraction based on multi-source data:

[0169] Construct a data association graph G(V,E), where the node set V contains work order data, feedback data, and public opinion data, and the edge set E represents the association between entities;

[0170] The association strength between nodes is calculated by weighted summation based on content similarity, temporal relevance, and spatial relevance, and an association strength matrix is ​​constructed.

[0171] Interaction features f are extracted based on the association strength matrix and data association graph. c =[f c1 ,f c2 ,f c3 ], where f c1 For relational features, f c2 For propagation characteristics, f c3 For association features;

[0172] In a specific example, the node set V contains three types of data nodes: Work order data nodes represent specific information about work orders (such as processing time, status, etc.). Feedback data nodes represent the content of user feedback (such as ratings, complaints, etc.). Public opinion data nodes represent the characteristics of public opinion texts or dissemination events. Each node is represented by its feature vector; for example, a work order node uses a time-series feature f. τ The feedback node uses semantic features f s Public opinion nodes use public opinion features f m Edges represent the relationships between nodes, and the edge weights are determined by the values ​​in the association strength matrix.

[0173] The elements w of the correlation strength matrix W ij Represents node v i and node v j The strength of the association between them is calculated using the following formula:

[0174] w ij =α·sim(c i ,c j )+β·exp(-|t i -t j | / τ t)+γ·exp(-d(p i ,p j ) / τ s )

[0175] Where: sim(c i ,c j ) is node v i and v j Content similarity, |t i -t j | represents the time difference between nodes, d(p) i ,p j ) represents the spatial distance between nodes, τ t ,τ s α, β, and γ are the time and space attenuation coefficients, and α, β, and γ are the weighting coefficients.

[0176] The association strength matrix W is a symmetric matrix that represents the association strength between all nodes.

[0177] Based on the data association graph G(V,E) and the association strength matrix W, the interaction features f are extracted. c =[f c1 ,f c2 ,f c3 The specific steps are as follows:

[0178] Relational feature f c1 The method for extracting information reflecting the network structure characteristics between nodes is as follows:

[0179] Calculate the degree centrality of nodes:

[0180]

[0181] It represents node v i The sum of the direct association strengths with other nodes.

[0182] Calculate the clustering coefficient of the nodes:

[0183]

[0184] Among them, L i For node v i The number of actual edges between neighbors, k i For node v i The degree of relation. Constructing relational features f c1 =[degree(v i ),C(v i )).

[0185] Propagation characteristics f c2 The extraction method reflects the ability of information to spread across the network and is as follows:

[0186] Calculate propagation efficiency:

[0187]

[0188] In the formula, d ij For node v i to node v j The shortest path length, where N is the total number of nodes in the network. Calculate the propagation range:

[0189]

[0190] In the formula, |V r | represents the number of reachable nodes, and |V| represents the total number of nodes in the network.

[0191] Calculate the propagation depth:

[0192]

[0193] In the formula, SPL is the shortest path length.

[0194] Constructing propagation features f c2 =[Pe,Ps,Pd].

[0195] Related feature f c3 The extraction method reflects the multi-dimensional relationships between nodes as follows:

[0196] Calculate content relevance:

[0197]

[0198] Calculate the time correlation:

[0199]

[0200] Calculate spatial correlation:

[0201]

[0202] Construct associated features f c3 =[Rc,Rt,Rs].

[0203] For time series features f τ Semantic features f s Public opinion characteristics f m Interaction features f c The system performs standardization and determines the fusion weights using an attention mechanism to obtain the final service quality feature F.

[0204] Specifically, in this embodiment, the risk assessment indicator system includes two levels of indicators. The first level of indicators includes service response indicators, service quality indicators, user satisfaction indicators, and social impact indicators. The second level of indicators are the next level of indicators set for the first level of indicators. Service response indicators include average response time and first response timeliness rate. Service quality indicators include problem resolution rate and service standardization. User satisfaction indicators include complaint rate and negative review rate. Social impact indicators include public opinion spread and negative impact degree.

[0205] The process of calculating power supply service risk assessment index values ​​based on service quality characteristics is as follows:

[0206] Construct a feature mapping function to calculate the second-layer index value vector I:

[0207] I=σ(W·F+b)·(1+tanh(F T ·M·F))

[0208] Where σ is the sigmoid function; F is the service quality feature; W is the feature mapping matrix; M is the feature interaction matrix; and b is the bias vector. Specifically, the feature mapping matrix W is an m×n matrix, where m is the number of second-layer indicators and n is the dimension of the service quality feature. It is obtained by training with historical data (supervised learning), initializing with expert experience, and then iteratively optimizing with data. The feature interaction matrix M is an n×n symmetric matrix, formed by calculating the mutual information between feature pairs.

[0209] This embodiment realizes a nonlinear mapping from the feature space to the index space, takes into account the interactive influence between features, ensures the range constraint of index values, and can adapt to different combinations of features.

[0210] Calculate the power supply service risk assessment index value R based on the second-level index value vector I:

[0211] R=ψ(I)·(1+η·‖ΔI‖)

[0212] In the formula, ψ(I) is a nonlinear combination function used to integrate the index values ​​of various dimensions, η is a time-varying parameter, ΔI is the change in the index value, and ‖·‖ represents the norm operation.

[0213] In this embodiment, ψ(I)=∑(w i ·I i )+∑(w ij ·I i ·I j ), w i The first-order weights reflect the importance of individual indicators, w ij The weights are second-order, reflecting the interactive importance between indicators, I i ,I jIt is the value of the second - layer index.

[0214] Specifically, in one embodiment of the present invention, the dynamic threshold model is established through the following steps:

[0215] Based on historical data, construct a time - series feature sequence, and use the long short - term memory network LSTM to extract the time - dependent feature z(t); z(t)=LSTM(F(t),h t-1 ), where F(t) is the quality - of - service feature at time t, and h t-1 is the hidden - layer state at the previous moment.

[0216] Calculate the adaptive factor λ(t):

[0217]

[0218] In the formula, u1 and u2 are adjustment parameters; ‖z(t)-z(t - 1)‖ represents the change amplitude of the features at the current moment and the previous moment.

[0219] Calculate the dynamic threshold boundaries T1(t) and T2(t), where T1(t)<T2(t):

[0220] T1(t)=μ his -λ(t)·σ his

[0221] T2(t)=μ his +λ(t)·σ his

[0222] In the formula, μ his is the mean value of the historical risk assessment index values, and σ his is the standard deviation of the historical risk assessment index values;

[0223] Determine the risk level according to the power supply service risk assessment index value R:

[0224] When R<T1(t), it is determined as a low - risk;

[0225] When T1(t)≤R<T2(t), it is determined as a medium - risk;

[0226] When R≥T2(t), it is determined as a high - risk.

[0227] The feature combination pattern represents the feature distribution law corresponding to different risk types, including the combination relationship and time - series evolution characteristics of each dimension feature in the quality - of - service feature F;

[0228] For each moment t in the historical data, construct a feature pattern matrix M(t):

[0229] M(t)=[F(t),ΔF(t),δF(t)]

[0230] Where F(t) is the historical service quality characteristic at time t, ΔF(t)=F(t)-F(t-1) is the change in the characteristic, and δF(t)=ΔF(t) / F(t-1) is the rate of change of the characteristic;

[0231] Based on a feature pattern matrix sequence {M(t)} constructed from historical data, a density clustering algorithm is used for analysis. Patterns with similar feature distributions and evolutionary patterns are clustered into one class, with each class corresponding to a specific risk type, thus forming a risk feature pattern library. In this embodiment, the density clustering algorithm used is DBSCAN, and the similarity between feature pattern matrices is calculated during clustering.

[0232]

[0233] In the formula, w k φ represents the feature importance weight. k (M i M j ) is the similarity measure of the k-th dimension feature. Based on the similarity, the feature pattern matrix is ​​clustered into several classes to form a risk feature pattern library.

[0234] For the current time t0, a feature pattern matrix M(t0) is constructed. The specific risk type is determined by calculating the similarity between M(t0) and each type of pattern in the risk feature pattern library. The similarity calculation employs a nonlinear measurement method that considers feature importance. The similarity calculation still uses the aforementioned similarity calculation formula to calculate S(M(t0), M... k M k This represents the k-th pattern in the risk feature pattern library. Based on the similarity S(M(t0), M... k The risk type corresponding to the pattern with the highest similarity is selected as the risk type at the current moment.

[0235] Specifically, in one embodiment of the present invention, the early warning output module receives the risk level and risk type, and forms the corresponding impact range and handling suggestions.

[0236] The scope of impact describes the geographical area, user group, or business scope that a risk event may affect. The scope of impact is calculated using an input data relationship diagram and interaction features f. c Propagating features f using interactive features c2 =[Pe,Ps,Pd], which determines the set of nodes V that may be affected by the risk by using the propagation range Ps and propagation depth Pd. r It also generates the scope of influence by combining node attributes (such as geographical location, user type, etc.).

[0237] The handling recommendations guide responses to risk events, including prioritization, resource allocation, and specific operational suggestions. First, handling recommendation templates are pre-defined for different risk types. For example: Service Quality Risk: Response Timeliness Risk: Recommend increasing customer service staff and optimizing response processes; Service Standardization Risk: Recommend strengthening employee training and improving service standards. User Satisfaction Risk: Complaint Escalation Risk: Recommend quickly handling user complaints and providing compensation solutions; User Churn Risk: Recommend conducting user follow-up activities and offering preferential policies. Social Impact Risk: Public Opinion Spread Risk: Recommend issuing an official statement to clarify the facts; Negative Impact Risk: Recommend communicating with the media to control the spread of public opinion. Operational Management Risk: Resource Allocation Risk: Recommend optimizing resource allocation and increasing backup resources; Process Execution Risk: Recommend checking process execution and promptly correcting problems. The priority of handling recommendations is adjusted according to the risk level: High Risk: Immediately activate the emergency plan and mobilize more resources; Medium Risk: Develop short-term response measures and closely monitor risk changes; Low Risk: Regularly check and track, and do a good job of prevention. Determine the resource allocation plan based on the scope of impact: if the scope of impact is large, it is recommended to mobilize more resources; if the scope of impact is concentrated in a certain region or user group, it is recommended to prioritize the handling of issues in that region or group.

[0238] Finally, the risk warning information generated by the warning output module includes the following: risk type: such as service quality risk, user satisfaction risk, etc.; risk level: low risk, medium risk, or high risk; scope of impact: including the affected geographical area, user group, or business scope; handling suggestions: specific response measures generated for the risk type and level.

[0239] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A power supply service risk identification system based on big data analysis, characterized in that, include: The data acquisition module is used to collect multi-source data related to power supply services, including work order data, feedback data, and public opinion data. The data processing module is used to preprocess multi-source data to obtain standardized data; The data fusion module is used to fuse standardized data to obtain fused data; The risk identification module is used to analyze the fused data, identify the types of risks in the power supply service, and determine the risk level by establishing a set of power supply service risk assessment indicators. The early warning output module is used to generate risk early warning information based on the risk identification results. The risk early warning information includes: risk type, risk level, scope of impact, and handling recommendations. A data registration algorithm is used to perform spatiotemporal alignment and feature fusion on multi-source data to achieve spatiotemporal consistency alignment of multi-source data. The calculation formula of the data registration algorithm is as follows: In the formula, This indicates the alignment result of position p at time t; This represents the observation value of the i-th data source at location p and time t; This represents the reference position p between position p and data source i. i Spatial distance between them; This represents the time t and the sampling time t of data source i. i The time difference between them; This represents the weight coefficient of the i-th data source; Represents the spatiotemporal resolution coefficient of data source i; Indicates the spatial resolution of data source i; Indicates the time resolution of data source i; , , These are adjustable parameters used to control spatial decay, temporal decay, and resolution effects. The risk identification module includes: The feature extraction unit is used to extract service quality features from the fused data; The risk assessment unit is used to construct a risk assessment indicator system and calculate the risk assessment indicator values ​​for power supply services based on service quality characteristics. The risk classification unit is used to establish a dynamic threshold model based on historical data statistics, determine the risk level according to the power supply service risk assessment index value and the dynamic threshold model, determine the risk type based on the feature combination pattern, and output the risk identification result. The execution process of the feature extraction unit is as follows: Extracting time-series features from work order data: The work order data is segmented using a sliding time window W(t,Δt), with the window length Δt adaptively adjusted according to the business characteristics. ,in This represents the average processing time for historical work orders. Standard deviation This is the lower limit of the window length; An LSTM network is constructed, incorporating attention-enhanced memory units. Information flow is controlled through forget gates, input gates, and output gates to extract temporal features. ,in, As an efficiency trend characteristic, through Due to the characteristics of business volume fluctuations, Features of time series patterns; Semantic feature extraction is performed on the feedback data: Construct a multi-level attention mechanism that includes word-level, sentence-level, and document-level attention: In the formula, is the word-level attention weight, with a value range of [0,1], and a sum of 1 within each sentence; represents the sentence-level attention weights, with a value range of [0,1] and a sum of 1 within the document; These are document-level attention weights, with a value range of [0,1] and a sum of 1. This is a word-level hidden layer representation, with the dimension being the word embedding dimension; This is a sentence-level hidden layer representation, with the dimension being the sentence encoding dimension; This is a document-level implicit representation, with the dimension being the document encoding dimension; , , This is the attention weight matrix for each layer; , , These are the attention vectors for each layer; Key information is identified by calculating the attention weights of each layer, and semantic features are extracted. ,in, Based on the theme distribution characteristics, Keyword features Emotional characteristics; Feature extraction from public opinion data: Constructing multimodal features of public opinion texts includes: extracting text features using a multi-level attention mechanism, extracting propagation network features through a graph convolutional network (GCN), and analyzing public opinion evolution using a time decay attention mechanism to extract temporal features; Calculate and analyze multimodal features to extract public opinion features. ,in, Characteristics of public opinion influence Characteristics of public opinion evolution This is a feature for public opinion clustering; Interaction feature extraction based on multi-source data: Construct a data association graph G(V,E), where the node set V contains work order data, feedback data, and public opinion data, and the edge set E represents the association between entities; The association strength between nodes is calculated by weighted summation based on content similarity, temporal relevance, and spatial relevance, and an association strength matrix is ​​constructed. Interaction features extracted based on association strength matrix and data association graph ,in, For relational features, As a characteristic of propagation, For association features; Temporal characteristics semantic features Public opinion characteristics Interaction features The system performs standardization and determines the fusion weights using an attention mechanism to obtain the final service quality feature F.

2. The power supply service risk identification system based on big data analysis as described in claim 1, characterized in that, The data processing module includes: The data cleaning unit is used to detect and process outliers in multi-source data. It uses an outlier detection algorithm based on local outlier factors to identify outlier data points and corrects outliers using a sliding window midpoint filling method. The data standardization unit is used to standardize the cleaned data. It adopts an adaptive standardization method based on the data distribution characteristics to transform data with different dimensions into a unified numerical range, thus obtaining standardized data.

3. The power supply service risk identification system based on big data analysis as described in claim 2, characterized in that, In the data standardization unit, the adaptive standardization method proceeds as follows: For a data sequence Y, calculate its distribution characteristic parameters α and β: In the formula, It is the upper quartile. is the lower quartile, and m is the sample size; The median. Standard deviation Skewness; Standardize the data sequence Y based on the distribution characteristic parameters α and β: In the formula, It is a smoothing factor used to prevent the denominator from being zero.

4. The power supply service risk identification system based on big data analysis as described in claim 1, characterized in that, The risk assessment indicator system includes two levels of indicators. The first level of indicators includes service response indicators, service quality indicators, user satisfaction indicators, and social impact indicators. The second - layer indicators are the next - level indicators corresponding to the first - layer indicators. The service response indicators include the average response duration and the timely first - response rate. The service quality indicators include the problem - solving rate and service standardization. The user satisfaction indicators include the complaint rate and the negative - review rate. The social influence indicators include the public - opinion diffusion degree and the negative - impact degree.

5. The power supply service risk identification system based on big data analysis as described in claim 4, characterized in that, The process of calculating the power - supply service risk - assessment indicator value based on service - quality characteristics is as follows: Construct a feature - mapping function to calculate the second - layer indicator - value vector I: in, It is the sigmoid function; For service quality characteristics; The feature mapping matrix; The feature interaction matrix; It is the bias vector; Calculate the power - supply service risk - assessment indicator value R based on the second - layer indicator - value vector I: In the formula, This is a non-linear combination function used to fuse index values ​​from various dimensions. For time-varying parameters, The change in the indicator value Represents norm operations.

6. The power supply service risk identification system based on big data analysis as described in claim 5, characterized in that, In the risk - classification unit, the dynamic - threshold model is established through the following steps: Construct a time - series feature sequence based on historical data, and use a long - short - term memory network (LSTM) to extract time - dependent features z(t); Calculate the adaptive factor λ(t): In the formula, , To adjust the parameters; Calculate the dynamic - threshold boundaries T1(t) and T2(t), where T1(t) < T2(t): In the formula, This represents the average of historical risk assessment indicator values. The standard deviation of historical risk assessment indicator values; Determine the risk level according to the power - supply service risk - assessment indicator value R: When R < T1(t), it is determined to be a low risk; When T1(t) ≤ R < T2(t), it is determined to be a medium risk; When R ≥ T2(t), it is determined to be a high risk.

7. The power supply service risk identification system based on big data analysis as described in claim 1, characterized in that, In the risk - classification unit, the feature - combination pattern represents the feature - distribution law corresponding to different risk types, including the combination relationship and time - series evolution characteristics of each dimension feature in the service - quality feature F; For each moment t in historical data, construct a feature - pattern matrix M(t): in, For the historical service quality characteristics at time t, The change in characteristics The rate of change of the characteristic; Based on the sequence of feature - pattern matrices {M(t)} constructed from historical data, use a density - clustering algorithm for analysis. Cluster the patterns with similar feature distributions and evolution laws into one category, and each category corresponds to a specific risk type, thus forming a risk - feature pattern library; Construct a feature - pattern matrix M(t0) for the current moment t0. Determine the specific risk type by calculating the similarity between M(t0) and each type of pattern in the risk - feature pattern library. The similarity calculation uses a non - linear measurement method considering feature importance.

Citation Information

Patent Citations

  • Power supply service risk early warning prevention and control platform and method

    CN114091889A