Accounting anomaly feature extraction method
Through multi-dimensional multi-index analysis and abnormal feature correlation diagram construction, abnormal features in power enterprise accounting business are extracted, and the stability and rapid response problems of power enterprise accounting business during business rules and electricity price adjustments are solved, achieving efficient abnormal detection and management.
Patent Information
- Application Number
- CN202111242327.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-25
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2041-10-25
AI Technical Summary
It is difficult for existing technology to respond quickly and ensure stability in the accounting business of power companies, especially when business rules and electricity prices are adjusted. How to scientifically detect and respond to emergencies is an urgent problem.
A method of calculating abnormal feature extraction is adopted to construct an exception library through multi-dimensional multi-index analysis and modeling, extract exception dimensions and abnormal index factors, build an abnormal feature correlation graph, and form a complete correlation graph structure through data fusion, deduce the correlation relationship between abnormal entities, and form an abnormal information correlation analysis model.
This method can accurately lock abnormal users, reduce abnormal situations in the accounting process caused by business changes, improve accounting management efficiency, and support the quality and efficiency improvement of related businesses in the power industry.
Smart Images

Figure CN113935819B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of information perception and recognition in the power industry, and relates to a method for extracting abnormal characteristics of accounting. Background Art
[0002] Accounting is the core business of the marketing system. Its main content is to accurately calculate the electricity bills and electricity consumption of each user. Statistical calculations are carried out according to relevant charging standards, and accounting treatments are made to record management materials.
[0003] With the rapid development of social construction, the demand for electric power energy is increasing day by day, resulting in a huge consumption of electric power energy. In this case, to effectively improve the economic benefits of power enterprises, higher requirements are put forward for accounting in power enterprise management. It is necessary not only to ensure the accuracy, authenticity and integrity of accounting data, but also to flexibly support local billing differences according to the differences between different power grids and provinces and policy situations, and support the expansion of dynamic and static business rules.
[0004] A large number of algorithm programs, rule changes, and electricity price adjustments have put forward higher requirements for the stability, accuracy and real-time performance of accounting. How to ensure the stability of the core business while making a quick response, and be able to conduct rapid detection based on scientific evidence for each change and adjustment, so that it has the ability to be put into production and the emergency response ability to handle emergencies, is an urgent problem to be solved at present. Summary of the Invention
[0005] To solve the deficiencies in the prior art, the present application provides a method for extracting abnormal characteristics of accounting. Considering the needs of the digital and intelligent transformation of the power industry, for the abnormal data generated in accounting, through multi-dimensional and multi-index analysis and modeling, an abnormal database is formed to support the accounting functions of different business scenarios and reduce the abnormal situations that occur in the accounting process due to business changes.
[0006] To achieve the above objectives, the present invention adopts the following technical solutions:
[0007] A method for extracting abnormal characteristics of accounting includes the following steps:
[0008] Step 1: According to the accounting business scenario, refine the abnormal data information of users, and integrate and standardize the abnormal data;
[0009] Step 2: After integrating and standardizing the abnormal data, construct a typical library of accounting anomalies;
[0010] Step 3: According to the typical library of accounting anomalies, analyze the abnormal influencing factors in multiple dimensions and multiple indicators, extract the abnormal dimensions and abnormal index factors, and construct an abnormal feature correlation diagram;
[0011] Step 4: Integrate the abnormal feature association graph with data, solve the association relationships between features from the abnormal feature association graph, fuse the feature data under multiple views to form a complete association graph structure, and deduce the association relationships between abnormal entities by analyzing the relationships between abnormal feature elements to obtain an abnormal information association analysis model;
[0012] Step 5: Train the abnormal information association analysis model, match abnormal entities through actual accounting business scenarios, and verify the relationship between abnormal entities and business scenarios.
[0013] The present invention further includes the following preferred solutions:
[0014] Preferably, the specific steps of Step 1 are as follows:
[0015] Step 1.1: Sort out the sources of abnormal data and obtain abnormal data information;
[0016] Step 1.2: Integrate the abnormal data information;
[0017] Step 1.3: Data standardization processing: including index unification processing and dimensionless processing.
[0018] Preferably, in Step 1.1, the data sources include accounting rules, quantity and fee refund and compensation processes, and business inspections. The corresponding abnormal data information is specifically:
[0019] Abnormal data generated according to accounting rules, and the information of clear abnormal data after analysis and judgment;
[0020] Data generated according to the quantity and fee refund and compensation process, combined with the reasons for refund and compensation and customer file information, and analyzed as the information of clear abnormal data;
[0021] Data sorted out and summarized according to daily business inspection work, and analyzed as the information of clear abnormal data.
[0022] Preferably, in Step 1.2, the abnormal data is integrated into a database through service interfaces, scheduling tasks, and ETL technology.
[0023] Preferably, in Step 1.3, the Z-score and Min-Max methods are used to standardize the data.
[0024] Preferably, the specific steps of Step 3 are as follows:
[0025] Step 3.1: According to the abnormal data information in the abnormal typical library, sort out and extract abnormal dimensions and abnormal index factors;
[0026] Step 3.2: Describe the abnormal entity as a set of multiple abnormal dimensions and abnormal index factors, perform feature extraction mapping on the abnormal entity, and construct an abnormal feature association graph of the abnormal entity.
[0027] Preferably, the abnormal dimensions include user classification, electricity consumption capacity, electricity consumption category, voltage level, industry classification, and time-sharing flag;
[0028] The abnormal index factors include total settlement electricity quantity, total settlement electricity charge, power factor adjustment electricity charge, and basic electricity charge.
[0029] Preferably, in step 3.2, the abnormal entity description is a set of abnormal dimension and abnormal index factor information:
[0030]
[0031] Among them, A i represents an abnormal entity, D represents a set of abnormal entities, f u represents a set of abnormal features, f represents an abnormal feature, m t represents the number of abnormal features of the abnormal entity A i under the t type;
[0032] For any abnormal entity A i ∈D, let F i be the feature set of this entity. After feature extraction mapping, the abnormal entity is added to the feature association graph, and the abnormal feature is mapped into the corresponding feature association graph, denoted as MAP(A i ) = {A i , F i}.
[0033] Preferably, the feature nodes have the following two types of association relationships:
[0034] (1): Correlation association relationship, which refers to the implicit relationship existing between abnormal features in the same view or different views;
[0035] (2): Indirect association relationship, which refers to the association relationship different from the correlation association existing between abnormal features in the same view or different views and needs to be deduced through the topological relationship of the feature association graph path.
[0036] Preferably, for the correlation association relationship, the correlation association relationship between two features is measured by entropy and mutual information;
[0037] For any nodes v i , v j with features f1, f2 ∈ R, R represents the set of abnormal features, I represents the mutual information of (f1, f2),
[0038] x and y represent random variables, p(x,y) represents the probability distribution for both x and y, p(x) represents the probability distribution for x, and p(y) represents the probability distribution for y;
[0039] Given a threshold δ, when I(f1,f2)>δ, it is considered that there is a correlation relationship between features f1 and f2, that is, there is a correlation relationship between nodes v i ,v j There is a correlation relationship between them.
[0040] Preferably, for an indirect association relationship, assume f1,f2,f3∈R. If Then there is That is, for features f1, f2, and f3 belonging to the same set R, if f1 and f2 are necessary and sufficient conditions for each other, and f2 and f3 are necessary and sufficient conditions for each other, then f1 and f3 are also necessary and sufficient conditions for each other;
[0041] Given a path e, it can include the threshold δ of e i . When e<δ, there is an indirect association relationship between non-adjacent nodes passed by the belonging path.
[0042] Preferably, in step 4, assume that the abnormal data set contains F types of feature elements. Let the abnormal feature element f i belong to the view φ(i), N represents the number of feature elements, and M represents the number of feature element views;
[0043] For two features f i ,f j in the abnormal feature space, construct the association relationship R ij between them in the complete feature space;
[0044] The association relationship between any two abnormal features can be divided into the same view association relationship φ(i)=φ(j) and the different view association relationship φ(i)≠φ(j);
[0045] s ij represents the association strength between the abnormal feature elements f i and f j . α φ(i)φ(j) represents the relative weight of the views φ(i)φ(j) in the abnormal information association analysis model space;
[0046] w i =[w i,1 ,w i,2 ,w i,3 ,…,w i,n , representing the association relationship expression of the feature element f i in the archive information association model space. Then there is w i =∑ 1<=j<=Nα φ(i)φ(j) , where ∑ 1<=j<=M a ij = 1;
[0047] For each feature element f i , 1 ≤ j ≤ N, w = [w1, w2, w3, …, w n T , there is w = R · w;
[0048]
[0049] Through this transformation, the feature elements of the partial view are transformed into feature elements in the abnormal data set space, and each abnormality is represented as a combination of multiple features in the space;
[0050] S(A i , A j ) represents the association relationship between the abnormal entities A i , A i ), then there is By solving the association strength between abnormal entities, a complete association graph structure is constructed to obtain an abnormal information association analysis model.
[0051] Preferably, in step 5, the business scenario is matched and accounted for. According to the user's business scenario, the user profile data is loaded, and the abnormalities that may be caused by the user in this scenario are simulated through the abnormal information association analysis model to verify the relationship between the abnormal entity and the business scenario, thereby reducing the abnormalities that occur during the accounting process due to business changes.
[0052] The beneficial effects achieved by this application:
[0053] The present invention integrates technologies such as data caching, in-memory computing, and cluster services, and uses intelligent analysis to conduct in-depth mining and analysis of accounting abnormalities. It can analyze the abnormalities generated in different scenarios of the accounting business in the power industry, extract abnormal features, accurately lock abnormal users, reduce the abnormalities that occur during the accounting process due to business changes, improve the management efficiency of accounting, and support the improvement of quality and efficiency of related power industry businesses. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 is a flowchart of a method for extracting accounting anomaly features of the present invention;
[0055] Figure 2 is the data processing and accounting anomaly typical library construction process in an embodiment of the present invention;
[0056] Figure 3 is a schematic diagram of abnormal feature extraction in an embodiment of the present invention;
[0057] Figure 4 It is a feature association diagram that is interconnected under multiple views in an embodiment of the present invention;
[0058] Figure 5 It is a model training flow chart in an embodiment of the present invention;
[0059] Figure 6 It is an abnormal feature extraction and application flow chart in an embodiment of the present invention. Detailed implementation manners
[0060] The present application will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and cannot be used to limit the protection scope of the present application.
[0061] As Figure 1 shown, a method for extracting accounting abnormal features of the present invention includes the following steps:
[0062] Step 1: According to the accounting business scenario, refine the abnormal data information of the user, and integrate and standardize the abnormal data. The specific steps are as follows:
[0063] Step 1.1: Sort out the sources of abnormal data and obtain abnormal data information;
[0064] In specific implementation, the data sources include accounting rules, quantity and fee refund and compensation processes, and business inspections. The corresponding abnormal data information is specifically:
[0065] The abnormal data generated according to the accounting rules, after analysis and judgment, becomes the information of clear abnormal data;
[0066] The data generated according to the quantity and fee refund and compensation process, combined with the reasons for refund and compensation and customer file information, is analyzed as the information of clear abnormal data;
[0067] The data sorted out and summarized according to the daily business inspection work is analyzed as the information of clear abnormal data.
[0068] Step 1.2: Integrate the abnormal data information: Integrate the abnormal data into a database through technologies such as service interfaces, scheduling tasks, and ETL.
[0069] Step 1.3: Data standardization processing: including index unification processing and dimensionless processing;
[0070] The index unification processing solves the problem of different natures between data;
[0071] The dimensionless processing solves the problem of comparability between data.
[0072] To generate different features, the Z-score and Min-Max methods are selectively used to standardize and normalize the data. Specifically:
[0073] Zero-mean normalization (z-score standardization)
[0074] Zero-mean normalization is also known as standard deviation standardization. After processing, the mean of the data is 0 and the standard deviation is 1.
[0075] The conversion formula is as follows:
[0076]
[0077] where, is the mean of the original data, σ is the standard deviation of the original data, and it is the most commonly used data standardization method currently. The standard deviation score can answer the question: "How many standard deviations is the given data from its mean?" Data above the mean will get a positive standardized score, and vice versa will get a negative standardized score.
[0078] Min-max normalization (Min-Max standardization)
[0079] Min-max normalization is also called discrete standardization. It is a linear transformation of the original data, mapping the data values to the range [0, 1].
[0080] The conversion formula is as follows:
[0081]
[0082] Deviation standardization retains the relationships existing in the original data and is the simplest method to eliminate the influence of dimension and data value range. The disadvantage of this processing method is that if the values are concentrated and a certain value is very large, then after normalization, each value is close to 0 and will not differ much.
[0083] For example, for the set of data 1, 1.2, 1.3, 1.4, 1.5, 1.6, 8.4. If in the future it encounters values outside the current [min, max] value range of the attribute, it will cause the system to report an error and the min and max need to be re-determined.
[0084] Step 2: After integrating and standardizing the abnormal data, construct a typical library for accounting anomalies to extract abnormal features. Steps 1-2 are as Figure 2 shown.
[0085] Step 3: Based on the typical library for accounting anomalies, analyze the abnormal influencing factors from multiple dimensions and multiple indicators, extract the abnormal dimension and abnormal indicator factors, and construct an abnormal feature correlation graph. The specific steps are as follows:
[0086] Step 3.1: According to the abnormal data information in the typical library for accounting anomalies, sort out and extract the abnormal dimension and abnormal indicator factors. The schematic diagram for abnormal feature extraction is as Figure 3 shown.
[0087] Furthermore, the abnormal feature extraction mainly includes four steps: standardization, normalization, feature selection, and chi-square selection.
[0088] Standardization means that for the samples in the training set, the data is divided by the variance or (and) the data is subtracted from its mean based on the column statistical information (the result is that the variance is equal to 1 and the data is near 0).
[0089] For example, when all features have a variance of 1 and / or a mean of 0, the radial basis function (RBF) kernel of SVM or the L1 and L2 regularized linear models usually have better effects.
[0090] Standardization can improve the convergence speed in the model optimization stage and also avoid the excessive influence of features with large variances on model training.
[0091] Normalization means that each independent sample is scaled so that the sample has a unit Lp norm. This is a common operation in text classification and clustering. For example, the dot product of two TF-IDF vectors that have been L2-normalized is the cosine similarity of the two vectors.
[0092] Feature selection means selecting the most relevant features for the modeling process. Feature selection reduces the size of the vector space, thereby reducing the time complexity of subsequent vector operations. The number of selected features can be adjusted through the validation set.
[0093] Chi-square selection means using the chi-squared method for feature selection. This method operates on labeled categorical data. Chi-square selection ranks the data based on the chi-square test and then selects the features with larger chi-square values (that is, the features most relevant to the label).
[0094] Extract the abnormal dimension and abnormal index factor. In order to represent the relationship between the abnormal entity A and the feature relationship f in the form of RDF triples, after the extraction of the abnormal dimension and abnormal index factor, the abnormal entity is expressed as a set of multiple abnormal dimensions and abnormal index factors.
[0095] The abnormal dimensions include user classification, electricity consumption capacity, electricity consumption category, voltage level, industry classification, time-sharing flag;
[0096] The abnormal index factors include total settlement electricity, total settlement electricity charge, power factor adjustment electricity charge, basic electricity charge, etc.
[0097] For example: Regarding the anomaly of "power factor assessment but no power factor adjustment charge", there is a certain high-voltage user. The power factor assessment method of this user's pricing strategy is standard assessment. The calculation method for the main metering point to participate in the power factor calculation is that both electricity quantity and electricity charge are involved. The calculation method for the sub-metering point to participate in the power factor calculation is that both electricity quantity and electricity charge are involved. However, there is no power factor adjustment charge, and the system prompts the anomaly of "power factor assessment but no power factor adjustment charge".
[0098] In this case, the anomaly dimensions are: customer classification, power factor assessment method of pricing strategy, calculation method for the main metering point to participate in the power factor calculation, calculation method for the sub-metering point to participate in the power factor calculation;
[0099] The anomaly index factor is: power factor adjustment charge, that is, the electricity charge for power factor adjustment.
[0100] The anomaly entity of "power factor assessment but no power factor adjustment charge" is jointly composed of the dimensions of customer classification, power factor assessment method of pricing strategy, calculation method for the main metering point to participate in the power factor calculation, calculation method for the sub-metering point to participate in the power factor calculation, and the power factor adjustment charge index factor.
[0101] Step 3.2: Describe the anomaly entity as a set of multiple anomaly dimensions and anomaly index factors, perform feature extraction mapping on the anomaly entity, and construct an anomaly feature association graph for the anomaly entity.
[0102] For example: Regarding the anomaly of "power factor assessment but no power factor adjustment charge", in the above Step 3.1, the anomaly entity is "power factor assessment but no power factor adjustment charge";
[0103] The anomaly dimensions are: customer classification, power factor assessment method of pricing strategy, calculation method for the main metering point to participate in the power factor calculation, calculation method for the sub-metering point to participate in the power factor calculation;
[0104] The anomaly index factor is: power factor adjustment charge, that is, the electricity charge for power factor adjustment.
[0105] After performing feature extraction on the anomaly entity of "power factor assessment but no power factor adjustment charge", its anomaly features are: the customer classification is a high-voltage user, the power factor assessment method of the pricing strategy is "standard assessment", the calculation method for the main metering point to participate in the power factor calculation is that both electricity quantity and electricity charge are involved, the calculation method for the sub-metering point to participate in the power factor calculation is that both electricity quantity and electricity charge are involved, and the power factor adjustment charge is 0. Thus, an anomaly feature association graph for the anomaly entity is constructed through these five anomaly features.
[0106] For any anomaly entity, after the anomaly feature extraction mapping step, the anomaly entity is added to the anomaly feature association graph.
[0107] The anomaly entity is described as a set of anomaly dimension and anomaly index factor information:
[0108]
[0109] Among them, A i represents an abnormal entity, D represents a set of abnormal entities, and f u represents a set of abnormal features, f represents an abnormal feature, and m t represents the number of abnormal features of the abnormal entity A i under the t type;
[0110] For any abnormal entity A i ∈D, let F i be the feature set of this entity. After feature extraction and mapping, the abnormal entity is added to the feature association graph, and the abnormal feature is mapped into the corresponding feature association graph, denoted as MAP(A i ) = {A i , F i}.
[0111] Due to the diversity of abnormal features, the feature nodes v i , v j ∈G i can be divided into the following two types of association relationships:
[0112] (1): Correlation association relationship, which refers to the implicit relationships such as dependence, restriction, and causality existing between abnormal features in the same view or different views.
[0113] Constructing the correlation association relationship means the process of analyzing the existing abnormal correlation relationships and finding the laws and patterns of the simultaneous occurrence of abnormal features based on statistical analysis.
[0114] The correlation association relationship between two features is mostly measured by entropy and mutual information;
[0115] For any nodes v i , v j with features f1, f2 ∈ R, where R represents the set of abnormal features, I represents the mutual information of (f1, f2),
[0116] x, y represent random variables, p(x, y) represents the probability distribution for both x and y, p(x) represents the probability distribution for x, and p(y) represents the probability distribution for y;
[0117] Given a threshold δ, when I(f1, f2) > δ, it is considered that there is a correlation relationship between features f1 and f2, that is, there is a correlation association relationship between nodes v i , v j .
[0118] (2): Indirect association relationship, which refers to the association relationship existing between abnormal features in the same view or different views that is different from the correlation association and needs to be deduced through the topological relationship of the feature association graph path.
[0119] Without loss of generality, assume that \(f_1, f_2, f_3\in R\). If then there is That is, for the features \(f_1\), \(f_2\), \(f_3\) belonging to the same set \(R\), if \(f_1\) and \(f_2\) are necessary and sufficient conditions for each other, and \(f_2\) and \(f_3\) are necessary and sufficient conditions for each other, then \(f_1\) and \(f_3\) are also necessary and sufficient conditions for each other;
[0120] The given path \(e\) can include the threshold \(\delta\) of \(e\) i When \(e < \delta\), there is an indirect correlation relationship between non - adjacent nodes passed by the belonging path.
[0121] The above two relationships aggregate abnormal data features of different types and attributes together to construct a feature correlation graph that is interconnected under multiple views, as Figure 4 shown.
[0122] For example: For the abnormality of "power factor assessment but no power adjustment charge", in the above step 3.2, the abnormal entity "power factor assessment but no power adjustment charge" contains five abnormal features:
[0123] The customer classification is high - voltage users, the pricing strategy power factor assessment method is "standard assessment", the participation method of the main metering point in power factor calculation is that electricity quantity participates and electricity charge participates, the participation method of the sub - metering point in power factor calculation is that electricity quantity participates and electricity charge participates, and the power adjustment charge is 0.
[0124] Among these five features, there is a correlation relationship between the participation method of the sub - metering point in power factor calculation and the participation method of the main metering point in power factor calculation. If the participation method of the main metering point in power factor calculation is that electricity quantity participates and electricity charge participates, then the participation method of the sub - metering point in power factor calculation must be that electricity quantity participates and electricity charge participates or electricity quantity participates and electricity charge does not participate; conversely, if the participation method of the sub - metering point in power factor calculation is that electricity quantity participates and electricity charge participates, then the participation method of the main metering point in power factor calculation must be that electricity quantity participates and electricity charge participates or electricity quantity participates and electricity charge does not participate. And there is an indirect correlation relationship between the pricing strategy power factor assessment method and the power adjustment charge. When the pricing strategy power factor assessment method is "standard assessment", there must be at least one metering point whose participation method in power factor calculation is that electricity quantity participates and electricity charge participates among the relevant metering points. When there is a metering point whose participation method in power factor calculation is that electricity quantity participates and electricity charge participates, the power adjustment charge can be calculated. Therefore, there is an indirect correlation relationship between the pricing strategy power factor assessment method and the power adjustment charge.
[0125] Step 4: Integrate the abnormal feature association graph with data, solve the association relationships between features from the abnormal feature association graph, and fuse the feature data under multiple views to form a complete association graph structure. Since each abnormal entity can be decomposed and mapped into multiple types of abnormal features, the association relationships between abnormal entities can be deduced by analyzing the relationships between abnormal feature elements, and an abnormal information association analysis model is obtained;
[0126] For example: For a high-voltage user with a single-rate pricing strategy type, the basic electricity charge calculation method is based on capacity, and there is no standby power supply. The system prompts an abnormality of "inconsistent basic electricity charge calculation method for single-rate users". There is a correlation association relationship between the pricing strategy type and the basic electricity charge calculation method. When the pricing strategy type is single-rate, the basic electricity charge calculation method can only be selected as not calculated, otherwise an abnormality of "inconsistent basic electricity charge calculation method for single-rate users" will occur.
[0127] Assume that the abnormal data set contains F types of feature elements, and let the abnormal feature element f i belong to the view φ(i), N represents the number of feature elements, and M represents the number of feature element views;
[0128] For two features f i , f j in the abnormal feature space, construct the association relationship R ij between them in the complete feature space;
[0129] The association relationship between any two abnormal features can be divided into the same-view association relationship φ(i) = φ(j) and the different-view association relationship φ(i) ≠ φ(j);
[0130] s ij represents the association strength between the abnormal feature elements f i and f j , and α φ(i)φ(j) represents the relative weight of the views φ(i)φ(j) in the abnormal information association analysis model space;
[0131] w i = [w i,1 , w i,2 , w i,3 , …, w i,n , representing the association relationship expression of the feature element f i in the archive information association model space. Then there is w i = ∑ 1<=j<=N α φ(i)φ(j) , where ∑ 1<=j<=M a ij = 1;
[0132] For each feature element f i, 1 ≤ j ≤ N, w = [w1, w2, w3, …, w n T , there is w = R·w;
[0133]
[0134] Through this transformation, the characteristic elements of the sub - views are transformed into characteristic elements in the abnormal data set space, and each abnormality is represented as a combination of multiple characteristics in the space;
[0135] S(A i , A j ) represents the association relationship between abnormal entities A i , A i . Then there is Construct a complete association graph structure by solving the association strength between abnormal entities to obtain an abnormal information association analysis model.
[0136] For example: For the two abnormal entities of "power factor assessment but no power adjustment fee" and "the metering point with residential electricity price wrongly implements power factor assessment", they are relevant under certain conditions.
[0137] There is a certain high - voltage user. The sub - metering point implements the combined electricity price for urban and rural residents. The power factor assessment method of the user pricing strategy is standard assessment. The power factor calculation method for the parent metering point is that both electricity quantity and electricity fee are involved. The power factor calculation method for the sub - metering point is that both electricity quantity and electricity fee are involved. The system prompts the abnormality of "the metering point with residential electricity price wrongly implements power factor assessment".
[0138] For the high - voltage user in step 3.2 above, there are the same abnormal characteristics: the customer classification is high - voltage user, the power factor assessment method of the pricing strategy is "standard assessment", the power factor calculation method for the parent metering point is that both electricity quantity and electricity fee are involved, and the power factor calculation method for the sub - metering point is that both electricity quantity and electricity fee are involved.
[0139] It shows that the two abnormal entities of "the metering point with residential electricity price wrongly implements power factor assessment" and "power factor assessment but no power adjustment fee" are relevant under certain conditions. At this time, an association graph between the abnormal entities of "the metering point with residential electricity price wrongly implements power factor assessment" and "power factor assessment but no power adjustment fee" can be constructed, and thus an abnormal information association analysis model for users with power factor assessment in the pricing strategy can be established.
[0140] Step 5: Train the abnormal information association analysis model, and match the abnormal entities through the actual accounting business scenario to verify the relationship between the abnormal entities and the business scenario.
[0141] In machine learning, training a model means learning (determining) the optimal values of all weights and biases using labeled samples. What machine learning algorithms do during the training process is to examine multiple samples and try to find a model that can minimize the loss to the greatest extent; the goal is to minimize the loss.
[0142] Furthermore, the model training process is as Figure 5 shown. The model training process mainly includes three links: the model, calculating the loss, and calculating the parameter update.
[0143] Model: Take one or more features as input and then return a prediction (y’) as output. For the sake of simplicity, consider a model that takes one feature and returns one prediction. The following formula (where b is the bias and w is the weight)
[0144] y′ = b + w1x1
[0145] Calculating the loss: Calculate the loss under these parameter (bias, weight) values through the loss function. In the sample space there are measurable states θ ∈ Θ and a random variable X, and the decision made according to the rule At this time, if there is a function L(θ, d) on the product space that satisfies: That is, for any L(θ, d), which is a non - negative measurable function, then L(θ, d) is called the loss function, representing the loss or risk corresponding to making decision d in state θ.
[0146] Calculating the parameter update: Detect the value of the loss function and generate new values for parameters such as bias and weight to minimize the loss.
[0147] For example: For the anomaly of "power factor assessment but no power adjustment charge", take the anomaly features of high - voltage users in step 4 above: the customer is classified as a high - voltage user, the power factor assessment method of the pricing strategy is "standard assessment", the way the master metering point participates in the power factor calculation is that electricity quantity participates and electricity charge participates, and the way the sub - metering point participates in the power factor calculation is that electricity quantity participates and electricity charge participates, as input, return "power factor assessment but no power adjustment charge" as output, and calculate the loss of the input anomaly features through the loss function.
[0148] Detect the value of the loss function, re - assign the weights and biases of the input anomaly features, calculate the loss again until the value of the loss function drops to the lowest.
[0149] After the model training is completed, it is matched with the actual accounting business scenarios such as "high-voltage user power factor standard assessment, main metering point electricity quantity participating in electricity charges, and sub-metering point electricity quantity participating in electricity charges". That is, the actual power factor is calculated by the sum of the active and reactive electricity quantities of the main metering point and the sub-metering point, and the power factor adjustment electricity charges are obtained by multiplying the catalog electricity charges of each metering point by the power factor, so as to verify the relationship between the abnormal entity of "power factor assessment but no power factor adjustment electricity charge" and the accounting business scenarios of "high-voltage user power factor standard assessment, main metering point electricity quantity participating in electricity charges, and sub-metering point electricity quantity participating in electricity charges".
[0150] The process of Step 3-5 is as Figure 6 shown.
[0151] Finally, it is matched with the accounting business scenario. According to the user's business scenario, the user profile data is loaded, and the possible anomalies that may be caused by the user in this scenario are simulated through the anomaly information correlation analysis model, thereby reducing the anomalies that occur during the accounting process due to business changes.
[0152] The present invention can extract the abnormal characteristics of multiple electricity business scenarios such as electricity charges from the accounting anomalies by analyzing different business scenarios of users, and can significantly reduce the abnormal situations generated during the accounting process from multiple dimensions around multiple business indicators, reduce the workload of personnel inspection and discrimination, and improve work efficiency.
[0153] The applicant of the present invention has made a detailed description and illustration of the embodiments of the present invention in combination with the accompanying drawings of the specification. However, those skilled in the art should understand that the above embodiments are only the preferred implementation schemes of the present invention, and the detailed description is only to help readers better understand the spirit of the present invention, rather than a limitation on the protection scope of the present invention. On the contrary, any improvement or modification based on the spirit of the present invention should fall within the protection scope of the present invention.
Claims
1. A method for extracting abnormal features in accounting, characterized in that: The method includes the following steps: Step 1: According to the accounting business scenario, refine the abnormal data information of users, and integrate and standardize the abnormal data, including: sorting out the sources of abnormal data and obtaining abnormal data information, where the data sources include accounting rules, quantity and fee refund and compensation processes, and business inspections. The corresponding abnormal data information is specifically: the abnormal data generated according to accounting rules, and the information of clearly abnormal data after analysis and judgment; the data generated according to the quantity and fee refund and compensation process, combined with the reasons for refund and compensation and customer file information, and analyzed as the information of clearly abnormal data; the data sorted out and summarized according to the daily business inspection work, and analyzed as the information of clearly abnormal data. Step 2: After integrating and standardizing the abnormal data, construct an accounting abnormal typical library. Step 3: According to the accounting abnormal typical library, analyze the abnormal influencing factors from multiple dimensions and multiple indicators, extract the abnormal dimension and abnormal index factors, and construct an abnormal feature correlation graph. The specific steps are as follows: Step 3.1: According to the abnormal data information in the abnormal typical library, sort out and extract the abnormal dimension and abnormal index factors. The abnormal dimensions include user classification, electricity consumption capacity, electricity consumption category, voltage level, industry classification, and time-sharing flag; the abnormal index factors include total settlement electricity quantity, total settlement electricity fee, power factor adjustment electricity fee, and basic electricity fee. Step 3.2: Describe the abnormal entity as a set of multiple abnormal dimensions and abnormal index factors, perform feature extraction mapping on the abnormal entity, and construct an abnormal feature correlation graph of the abnormal entity. The abnormal entity is described as a set of abnormal dimension and abnormal index factor information: Among them, A i represents an abnormal entity, D represents a set of abnormal entities, and f u represents a set of abnormal features, f represents an abnormal feature, and m t represents the number of abnormal features of the abnormal entity A i under the t type; For any abnormal entity A i ∈ D, let F i be the feature set of this entity. After feature extraction mapping, the abnormal entity is added to the feature association graph, and the abnormal features are mapped into the corresponding feature association graph, denoted as MAP(A i ) = {A i , F i}; The feature nodes have the following 2 types of association relationships: (1): Correlation association relationship, which refers to the implicit relationship existing between abnormal features in the same view or different views. (2): Indirect association relationship, which refers to the association relationship different from the correlation association existing between abnormal features in the same view or different views, and needs to be deduced through the path topology relationship of the feature correlation graph. For the correlation association relationship, the correlation association relationship between two features is measured by entropy and mutual information. For any node v i , v j with features f1, f2 ∈ R, where R represents the set of abnormal features and I represents the mutual information of (f1, f2), where x, y represent random variables, p(x, y) represents the probability distribution for both x and y, p(x) represents the probability distribution for x, and p(y) represents the probability distribution for y; given a threshold δ, when I(f1, f2) > δ, it is considered that there is a correlation between features f1 and f2, that is, there is a correlation association between nodes v i , v j ; For an indirect association relationship, assume f1, f2, f3 ∈ R. If then there is e = e1° That is, for the features f1, f2, f3 belonging to the same set R, if f1 and f2 are necessary and sufficient conditions for each other, and f2 and f3 are necessary and sufficient conditions for each other, then f1 and f3 are also necessary and sufficient conditions for each other; the given path e can include the threshold δ of e i When e < δ, there is an indirect association relationship between the non-adjacent nodes passed by the path to which it belongs; Step 4: Integrate the abnormal feature correlation graph with data, solve the association relationship between features from the abnormal feature correlation graph, fuse the feature data under multiple views to form a complete association graph structure, and deduce the association relationship between abnormal entities by analyzing the relationship between abnormal feature elements to obtain an abnormal information association analysis model. Step 5: Train the abnormal information association analysis model, and match the abnormal entity through the actual accounting business scenario to verify the relationship between the abnormal entity and the business scenario.
2. A method for extracting abnormal features in accounting according to claim 1, characterized in that: The specific steps of Step 1 are: Step 1.1: Sort out the sources of abnormal data and obtain abnormal data information. Step 1.2: Integrate the abnormal data information. Step 1.3: Data standardization processing: including index unification processing and dimensionless processing.
3. A method for extracting abnormal features in accounting according to claim 2, characterized in that: In step 1.2, the abnormal data is integrated into a database through service interfaces, scheduling tasks, and ETL technology.
4. A method for extracting accounting anomaly features according to claim 2, characterized in that: In step 1.3, the Z-score and Min-Max methods are used to standardize the data.
5. A method for extracting accounting anomaly features according to claim 1, characterized in that: In step 4, assume that the abnormal dataset contains F types of feature elements, and let the abnormal feature element f i belong to the view φ(i), N represents the number of feature elements, and M represents the number of views of feature elements; For two features f i , f j in the abnormal feature space, construct the association relationship R ij ; The correlation relationship between any two anomaly features can be divided into the same-view correlation relationship φ(i) = φ(j) and the different-view correlation relationship φ(i) ≠ φ(j); s ij represents the abnormal feature element f i and f j the association strength of, α φ(i)φ(j) represents the relative weight of the view φ(i)φ(j) in the abnormal information association analysis model space; w i = [w i,1 , w i,2 , w i,3 , …, w i,n , representing the association relationship expression of the characteristic element f i in the file information association model space, then there is w i = ∑ 1<=j<=N α φ(i)φ(j) , where ∑ 1<=j<=M a ij = 1; For each feature element f i , 1 ≤ j ≤ N, w = [w1, w2, w3, …, w n T , there is w = R · w; Through this transformation, the feature elements of the sub-views are transformed into feature elements in the abnormal data set space, and each anomaly is represented as a combination of multiple features in the space; S(A i ,A j ) represents the abnormal entity A i ,A i For the association relationship between them, there is By solving the association strength between abnormal entities, a complete association graph structure is constructed to obtain an abnormal information association analysis model.
6. A method for extracting accounting anomaly features according to claim 1, characterized in that: In step 5, match the accounting business scenario, load the user profile data according to the user's business scenario, simulate the anomalies that may be caused by the user in this scenario through the anomaly information correlation analysis model, verify the relationship between the anomaly entity and the business scenario, and thus reduce the anomalies occurring in the accounting process due to business changes.
Citation Information
Patent Citations
Power accounting anomaly analysis method and system
CN113076353A
Abnormal information processing node analysis method and apparatus, medium and electronic device
WO2021159834A1