An Optimization Method for Power Big Data Management Combining Data Owners

By combining the data owner's power big data management optimization method, including data owner identification and the application of blockchain technology, data quality and attribution problems in power big data management are solved, data management automation and responsibility accuracy are realized, and data security and credibility are ensured.

CN119831183BActive Publication Date: 2025-06-24STATE GRID INFO TELECOM GREAT POWER SCI & TECH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510316277.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-06-24
Estimated Expiration
2045-03-18

AI Technical Summary

Technical Problem

In the process of collecting, storing, processing and circulation, power big data faces problems such as uneven data quality, fuzzy data belonging, and difficult to guarantee data authenticity, security and traceability.

Method used

A power big data management optimization method combined with data owners is adopted. By collecting power big data, extracting context features, identifying the initial master, building a data owner identification model, automatically writing to a data governance platform, recording the identification process using blockchain technology, and detecting data problems through rule base detection and abnormal detection models, generating work orders and pushing them to the data owner for repair.

Benefits of technology

It realizes the automation of data management, the accuracy of responsibility flow and the transparency of safety and trustworthiness, improves data quality control, solves the problem of complexity of ownership of power data, and ensures the authenticity, security and traceability of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119831183B_ABST
    Figure CN119831183B_ABST
Patent Text Reader

Abstract

The present invention relates to an optimization method for power big data management in combination with data owners, comprising the following steps: S1: Collect power big data and extract context features; S2: Identify the initial owner according to metadata ownership information or data uploader; S3: Construct a data owner identification model, and based on the extracted context features, achieve accurate identification and dynamic update of the data owner, correct the initial owner, and obtain the final data owner; S4: Automatically write the result output by the data owner identification model into the data governance platform to form a structured data ownership record, and use blockchain technology to record the identification process; S5: Use a rule library detection and anomaly detection model to detect data problems; and generate work orders for the detected data problems; S6: After receiving the problems, the data owner optimizes and repairs the problem data, and verifies whether the data repair is compliant. If it is not compliant, a reverse work order is returned for continued processing. The present invention realizes high-quality, high-efficiency and high-trust data management.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of big data management, and particularly to an optimization method for power big data management combined with data owners. Background Art

[0002] In recent years, with the digital transformation and intelligent development of the power industry, power big data has become an important foundation for supporting applications such as power grid dispatching optimization, load forecasting, fault diagnosis, and user behavior analysis. However, power big data faces many technical challenges in the process of collection, storage, processing, and circulation: First, the data sources are diverse (such as sensor data, dispatching logs, user power consumption behavior data, etc.), resulting in uneven data quality, with problems such as missing fields, duplicates, and outliers; Second, the data ownership is often ambiguous, such as it is difficult to accurately divide data responsibilities in a data sharing environment with multi-party collaboration; In addition, power data governance needs to ensure the authenticity, security, and traceability of data to comply with industry regulations and support multi-party collaboration. Summary of the Invention

[0003] In order to solve the above problems, the purpose of the present invention is to provide an optimization method for power big data management combined with data owners, realizing the automation of data management, the precision of responsibility transfer, and the transparency of security and trustworthiness.

[0004] To achieve the above purpose, the present invention adopts the following technical solutions:

[0005] An optimization method for power big data management combined with data owners, comprising the following steps:

[0006] S1: Collect power big data and extract context features;

[0007] S2: Identify the initial owner according to the metadata ownership information or the data uploader;

[0008] S3: Construct a data owner identification model, and based on the extracted context features, achieve the precise identification and dynamic update of the data owner, correct the initial owner, and obtain the final data owner;

[0009] S4: Automatically write the result output by the data owner identification model into the data governance platform to form a structured data ownership record, and use blockchain technology to record the identification process;

[0010] S5: Use the rule library detection and anomaly detection model to detect data problems; and generate work orders for the detected data problems, and immediately push the work orders according to the data owners identified by the data owner identification model;

[0011] S6: After receiving the question, the data owner optimizes and repairs the question data, and verifies whether the data repair is compliant. If it is not compliant, a reverse work order is returned for further processing.

[0012] Furthermore, the power big data includes power consumption behavior data, user operation behavior records, and metadata; the power consumption behavior data includes user power consumption P(t), power grid equipment operation data, time series type data, and special scenario records; the user operation behavior records are audit logs, including access, modification, and read frequency data; the metadata includes the table structure, data fields, and field description information of the power internal management system.

[0013] Furthermore, the context features include power consumption behavior pattern values, operation behavior characteristics, and metadata characteristics, which are specifically as follows:

[0014] The power consumption behavior pattern values include peak load periods and abnormal fluctuations. Calculate the daily load value from the user power consumption P(t), and extract the peak P peak and the corresponding time period t peak :

[0015] ;

[0016] Calculate the difference between the user power consumption at time t + 1 and the user power consumption at time t , if ΔP>T threshold , then the corresponding time t is marked as an abnormal fluctuation point, where T threshold is a preset threshold:

[0017]

[0018] where mean and std are the mean and standard deviation of the fluctuation, and k is an adjustment coefficient;

[0019] The operation behavior characteristics include the data usage frequency F access (u) and modification preferences :

[0020] ;

[0021] ;

[0022] where, N access (u) is the number of times user u accesses the data within time T; is the total number of modification operations for user u on field f i all, is the total number of modifications for user u on all fields;

[0023] The metadata features include the extraction of field semantic relevance and the naming patterns extracted from table names or field names; the field semantic relevance is calculated using cosine similarity:

[0024] ;

[0025] where V( f 1 ), V( f 2 ) are the embedding vectors of fields f 1 ,f 2 respectively; are the norms of V( f 1 ), V( f 2 ) respectively.

[0026] Furthermore, S2 is specifically as follows:

[0027] Collect the metadata ownership and field description information by querying the metadata table of the SQL query system; use the ELK log collection and processing framework to collect user operation logs in real time;

[0028] Initialize the unified field structure of log data and metadata, associate the table data by time and user ID, parse and classify the field names, and classify the field names into specific responsibility types;

[0029] Give priority to matching the owner_id information in the metadata. If there is no owner record, match the metadata field description information or the creator identity; if the metadata cannot be recognized, identify the initial owner based on the upload log and the most recent upload operation.

[0030] Furthermore, the construction of the data owner recognition model is specifically as follows:

[0031] Convert the context features into a structured vector for model input, construct the input feature sequence X, and each owner candidate corresponds to a context feature sequence, and represent each context feature sequence as a time-related multi-dimensional vector:

[0032] ;

[0033] where x t is the feature at the t-th moment, including the multi-dimensional context feature vector;

[0034] ;

[0035] where is the owner information in the metadata; is the field and responsibility matching score; is the data upload timestamp;

[0036] Use the embedding layer to encode x t into a fixed-length vector e t , then the input feature sequence X is transformed into the embedding feature sequence E:

[0037] ;

[0038] The input sequence is the embedding feature sequence E, and the time context information is added through the positional encoding P to obtain the final input Z = E + P;

[0039] Use the self-attention mechanism of the Transformer to model the mutual relationship between context features and calculate the weight at each time step :

[0040] ;

[0041] where Q = ZW Q is the query matrix, K = ZW K is the key matrix, V = ZW V is the value matrix; d k is the dimension of the query / key; G is the dense matrix; W Q、 W K、 W V are the weights of the query matrix, key matrix, and value matrix respectively;

[0042] Capture the dependencies between different features through the multi-head mechanism:

[0043] ;

[0044] where, is the weight coefficient; is the th attention head; represents the concatenation function; is the output of the multi-head mechanism;

[0045] ;

[0046] where, is the trainable matrix of the attention head ; After l layers of Transformer encoding, the output feature sequence is obtained ;

[0047] And use global pooling to aggregate the sequence information to obtain the comprehensive recognition vector H = Pool(Z l );

[0048] Map the output to the category space of all candidate responsible persons through a fully connected layer:

[0049] ;

[0050] Among them, and are the weight and bias of the fully connected layer respectively; is the probability that the responsible person is u;

[0051] Then, the finally identified responsible person is: 。

[0052] Furthermore, the dense matrix G is constructed based on GNN as follows:

[0053] Regard the input feature sequence X as the node representation of the graph, and each time-step feature x t corresponds to a node; the initial adjacency matrix A is a fully connected graph and is initialized as the identity matrix;

[0054] Define the node update formula in the m-th layer of GNN:

[0055] ;

[0056] Among them, is the hidden feature of node a in the m-th layer; is the hidden feature of node b in the m-th layer; represents the relationship strength between node a and node b; is the adjacent node of node a; 、 are the learnable weights and biases in GNN respectively;

[0057] The node hidden features will be updated in each layer, and at the same time, the adjacency matrix will be adjusted according to the feature changes to obtain a dynamically updated adjacency matrix:

[0058] ;

[0059] Among them, is the activation function; is the similarity function;

[0060] Finally, after M layers of GNN updates, the output adjacency matrix is the dynamically generated dense matrix 。

[0061] Furthermore, the S4 is specifically:

[0062] The data owner recognition model processes the input data, generates the ownership result, interacts with the data governance platform through the interface, and transmits the structured ownership data;

[0063] The data governance platform completes the data storage and permission allocation update, and at the same time triggers the blockchain writing process;

[0064] The blockchain records the recognition process and ownership changes, and supports traceability and historical record query;

[0065] And synchronize the blockchain with the data governance platform database to ensure that all ownership changes in the data governance platform are reflected on the chain, and verify the data consistency of the platform from the historical records on the chain.

[0066] Furthermore, S5 is specifically

[0067] Build a rule library, store predefined detection rules and meta-information, configure the anomaly detection model, and perform adaptive operation after parameterized training;

[0068] And based on the rule library and the anomaly detection model, detect data problems for timed or real-time scenarios;

[0069] Generate a structured problem record for each detected problem, and organize the detected data problems into a work order readable for the data owner;

[0070] Push the work order to the data owner responsible for repair or review through the determined data owner output by the data owner recognition model.

[0071] Furthermore, S6 is specifically:

[0072] The data owner receives the problem work order through the work order receiving platform, including the problem description, the impact scope, and the repair suggestion;

[0073] After the data owner repairs the data, submit the repair result to the data governance platform; the data governance platform calls the compliance verification module to verify the repair result:

[0074] If the repair is compliant, the work order is closed; if the repair is non-compliant, generate a reverse work order, send the problem and the verification feedback to the data owner, and require further processing.

[0075] The present invention has the following beneficial effects:

[0076] 1. The present invention realizes high-quality, high-efficiency and high-trust data management, can not only strengthen data quality control, but also improve the overall data usage value of the power industry, and lay a solid foundation for the efficient operation of the smart grid;

[0077] 2. The present invention dynamically constructs a data owner recognition model in combination with context features, more accurately divides data responsible persons, corrects from the initial owner to the final context-aware owner, effectively solves the problem of the complexity of power data ownership (such as the ambiguous attribution of multi-source data and shared data), and adopts a dynamic update mechanism to ensure real-time recognition and attribution adjustment during data generation, transfer, and use; in combination with blockchain technology, key information such as the data owner recognition process, data ownership records, and work orders is stored on the blockchain to ensure immutability and the credibility of the whole-process record;

[0078] 3. The present invention uses rule library detection (such as problems like missing fields and duplicate values) in combination with an anomaly detection model (such as distribution anomaly, cross-dimensional anomaly, and time-series anomaly) to achieve comprehensive coverage and intelligent detection of data problems, and based on the data owner recognition model combined with fast work order push, enables each problem to accurately find the responsible person, maximally shortening the problem response time and repair time. After the repair is completed, the compliance is automatically verified, and a benign transfer mechanism is established to reduce friction during the responsibility attribution and repair process. BRIEF DESCRIPTION OF THE DRAWINGS

[0079] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0080] The following further describes the present invention in detail with reference to the drawings and specific embodiments:

[0081] Reference Figure 1 , in this embodiment, an optimization method for power big data management in combination with data owners is provided, including the following steps:

[0082] S1: Collect power big data and extract context features;

[0083] S2: Identify the initial owner according to the metadata ownership information or the data uploader;

[0084] S3: Construct a data owner recognition model, based on the extracted context features, achieve accurate recognition and dynamic update of the data owner, correct the initial owner, and obtain the final data owner;

[0085] S4: Automatically write the result output by the data owner recognition model into the data governance platform to form a structured data ownership record, and use blockchain technology to record the recognition process to ensure immutability and achieve traceability;

[0086] S5: Use rule library detection (such as missing fields and duplicate values) and an anomaly detection model to detect data problems; and generate work orders for the detected data problems (the work order content includes problem description, scope of influence, etc.), and immediately push the work orders according to the data owner identified by the data owner recognition model;

[0087] S6: After receiving the question, the data owner optimizes and repairs the question data, and verifies whether the data repair is compliant. If it is not compliant, a reverse work order is returned for further processing.

[0088] In this embodiment, the power big data includes power consumption behavior data, user operation behavior records, and metadata; the power consumption behavior data includes user power consumption P(t), power grid equipment operation data (load power, abnormal fluctuation records), time series type data (such as voltage, current, power and other fluctuation data), and special scenario records (identifying faults, load peaks, power outage events, etc.); the user operation behavior records are audit logs, including access, modification, and read frequency data; the metadata includes the table structure, data fields, and field description information of the power internal management system.

[0089] In this embodiment, the context features include power consumption behavior pattern values (such as peak load periods, abnormal fluctuation characteristics), operation behavior characteristics (such as data usage frequency, modification preferences), and metadata characteristics (semantic associations, naming rules, etc.), which are specifically as follows:

[0090] The power consumption behavior pattern values include peak load periods and abnormal fluctuations. Calculate the daily load value from the user power consumption P(t), and extract the peak P peak and the corresponding period t peak :

[0091] ;

[0092] Calculate the difference between the user power consumption at time t + 1 and the user power consumption at time t , if ΔP>T threshold , then the corresponding time t is marked as an abnormal fluctuation point, where T threshold is a preset threshold:

[0093]

[0094] where mean and std are the fluctuation mean and standard deviation, and k is an adjustment coefficient;

[0095] The operation behavior characteristics include data usage frequency F access (u) and modification preferences :

[0096] ;

[0097] ;

[0098] where, N access (u) is the number of times user u accesses the data within time T; is for user u for field fi The total number of modification operations is the total number of modifications made by user u on all fields;

[0099] The metadata features include the extraction of field semantic relevance and the naming patterns extracted from table names or field names; the field semantic relevance is calculated using cosine similarity:

[0100] ;

[0101] where V( f 1 ), V( f 2 ) are the embedding vectors of fields f 1 ,f 2 respectively; are the norms of V( f 1 ), V( f 2 ) respectively.

[0102] In this embodiment, S2 is specifically as follows:

[0103] Query the metadata table of the system through SQL to collect metadata ownership and field description information; use the ELK log collection and processing framework to collect user operation logs in real time;

[0104] Initialize the unified field structure of log data and metadata, associate table data by time and user ID, parse and classify field names, and classify field names into specific responsibility types;

[0105] Give priority to matching the owner_id information (owner information) in the metadata. If there is no owner record, match the metadata field description information or the creator identity; if the metadata cannot be recognized, identify the initial owner based on the upload log and the most recent upload operation.

[0106] In this embodiment, the construction of the data owner identification model is specifically as follows:

[0107] Convert the context features into a structured vector for model input, construct the input feature sequence X, where each owner candidate corresponds to a context feature sequence, and represent each context feature sequence as a time-related multi-dimensional vector:

[0108] ;

[0109] where x t is the feature at the t-th moment, including a multi-dimensional context feature vector;

[0110] ;

[0111] Among them, is the owner information in the metadata; is the field and responsibility matching score; is the data upload timestamp;

[0112] Use the embedding layer to encode x t into a fixed-length vector e t , then the input feature sequence X is transformed into the embedding feature sequence E:

[0113] ;

[0114] The input sequence is the embedding feature sequence E, and the time context information is added through the positional encoding P to obtain the final input Z = E + P;

[0115] Use the self-attention mechanism of the Transformer to model the mutual relationship between context features and calculate the weight at each time step :

[0116] ;

[0117] Among them, Q = ZW Q is the query matrix, K = ZW K is the key matrix, V = ZW V is the value matrix; d k is the dimension of the query / key; G is the dense matrix; W Q、 W K、 W V are the weights of the query matrix, key matrix, and value matrix respectively;

[0118] Capture the dependencies between different features through the multi-head mechanism:

[0119] ;

[0120] Among them, is the weight coefficient; is the th attention head; represents the concatenation function; is the output of the multi-head mechanism;

[0121] ;

[0122] Among them, is the trainable matrix of the attention head ;

[0123] After l - layer Transformer encoding, the output feature sequence is obtained ;

[0124] And use global pooling Pool to aggregate sequence information to obtain the comprehensive recognition vector H = Pool(Z l );

[0125] Through the fully - connected layer Map the output to the class space of all candidate responsible parties:

[0126] ;

[0127] Among them, and are the weight and bias of the fully - connected layer respectively; is the probability that the responsible party is u;

[0128] Then, the final recognized responsible party is: .

[0129] In this embodiment, the dense matrix G is constructed based on GNN as follows:

[0130] Regard the input feature sequence X as the node representation of the graph, and each time - step feature x t corresponds to a node; the initial adjacency matrix A is a fully - connected graph and is initialized as the identity matrix;

[0131] Define the node update formula in the m - th layer of GNN:

[0132] ;

[0133] Among them, is the hidden feature of node a in the m - th layer; is the hidden feature of node b in the m - th layer; represents the relationship strength between node a and node b; is the adjacent node of node a; 、 are the learnable weight and bias in GNN respectively;

[0134] The node hidden feature will be updated in each layer, and at the same time, the adjacency matrix is adjusted according to the feature change to obtain the dynamically updated adjacency matrix:

[0135] ;

[0136] Among them, is the activation function; is the similarity function;

[0137] Finally, after M layers of GNN updates, the output adjacency matrix is the dynamically generated dense matrix 。

[0138] In this embodiment, S4 is specifically as follows:

[0139] The data owner recognition model processes the input data, generates the ownership result, interacts with the data governance platform through the interface, and transmits the structured ownership data;

[0140] The data governance platform completes data storage and permission allocation updates, and at the same time triggers the blockchain writing process;

[0141] The blockchain records the recognition process and ownership changes, supporting traceability and historical record query;

[0142] And synchronize the blockchain with the data governance platform database to ensure that all ownership changes in the data governance platform are reflected on the chain, and verify the data consistency of the platform from the historical records on the chain.

[0143] In this embodiment, S5 is specifically

[0144] Build a rule library, store predefined detection rules and meta-information (such as field constraints, data distribution), and configure an anomaly detection model to perform adaptive operation after parameterized training;

[0145] The rule library uses predefined rules to implement field-level or table-level problem detection based on specific business requirements and data governance specifications.

[0146] Field missing: Check whether the corresponding field in the table is empty or some field values are not filled.

[0147] Duplicate values: Check whether there are primary key conflicts or duplicate records in the table.

[0148] Data constraint verification: Check whether the field conforms to the defined range (such as date field format, numerical range, character length limit).

[0149] Business logic verification: Define rules based on specific business logic (such as transaction amount must >0, time sequence must increase).

[0150] Rule implementation method:

[0151] According to the business-defined rule template (rule name, target field, detection conditions, etc.).

[0152] Detect directly based on SQL query or implement through scripts;

[0153] For problems that cannot be pre-set with rules and require dynamic judgment (such as outliers, distribution changes, etc.), an anomaly detection model is used, and the anomaly detection model is constructed based on One-Class SVM.

[0154] Based on the rule library and the anomaly detection model, for scenarios such as timing (e.g., daily / hourly) or real-time (when data is updated), detect data problems;

[0155] Generate a structured problem record for each detected problem, and organize the detected data problems into a work order that is readable to the data owner;

[0156] Preferably, the content of the work order includes:

[0157] Basic information;

[0158] Work order number: The unique identifier of the work order;

[0159] Submission time, submitter: The time when the work order is generated and the generation system;

[0160] Data information: The ID of the problem data, metadata information (such as generation time, data source, table name, etc.);

[0161] Problem details;

[0162] Problem type: Field missing, duplicate value, outlier, etc.;

[0163] Problem description: The specific description of the detected problem (such as "The number of missing values in the field exceeds 120 rows, and the coverage ratio is 25%");

[0164] Problem impact: Clearly indicate which business logics or data links may be affected by this problem;

[0165] Impact scope: The specific data and fields affected;

[0166] Fixing solution;

[0167] Suggested data repair or remedial solution.

[0168] Push the work order to the data owner for repair or review through the determined data owner output by the data owner identification model.

[0169] The push method is as follows:

[0170] Obtain data owner information;

[0171] Call the data owner identification model and query the unique identifier of the responsible person according to the data ID;

[0172] The model returns the data owner Owner ID and other meta-information (such as the scope of authority) to determine the specific person in charge.

[0173] Push work order content:

[0174] Online instant push:

[0175] Directly push the work order to the user notification interface of the data person in charge through the message system of the data governance platform;

[0176] External system integration:

[0177] Pass the work order to the corresponding business system through REST API or message queue (such as Kafka, RabbitMQ, etc.).

[0178] Notify the person in charge of relevant work order tasks by email or text message.

[0179] In this embodiment, S6 specifically is:

[0180] The data owner receives the problem work order through the work order receiving platform, including problem description, impact scope, and repair suggestions;

[0181] After the data owner repairs the data, submit the repair result to the data governance platform; the data governance platform calls the compliance verification module to verify the repair result:

[0182] If the repair is compliant, the work order is closed; if the repair is non-compliant, generate a reverse work order, send the problem and verification feedback to the data owner, and require further processing.

[0183] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0184] The present invention is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of the flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate for implementing in the process Figure 1 one process or multiple processes and / or blocks Figure 1a device for the functions specified in one or more boxes.

[0185] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device that implements the functions specified in one Figure 1 process or more processes and / or boxes Figure 1 a box or more boxes.

[0186] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one Figure 1 process or more processes and / or boxes Figure 1 a box or more boxes.

[0187] As described above, it is only the preferred embodiment of the present invention, and it is not intended to limit the present invention in other forms. Any person skilled in the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A method for optimizing power big data management in combination with data owners, characterized in that: The following steps are involved: S1: Collect power big data and extract context features; S2: Identify the initial owner based on metadata ownership information or data uploader; S3: Build a data owner identification model to accurately identify and dynamically update the data owner based on the extracted context features, modify the initial owner, and obtain the final data owner; S4: Automatically write the output of the data owner identification model into the data governance platform to form a structured data ownership record, and use blockchain technology to record the identification process; S5: Use the rule base detection and anomaly detection model to detect data problems; generate work orders for the detected data problems, and push the work orders immediately based on the data owner identified by the data owner identification model; S6: After receiving the problem, the data owner optimizes and repairs the problematic data and verifies whether the data repair is compliant. If not, a reverse work order is returned for further processing; The construction of the data owner identification model is as follows: The context features are converted into structured vectors for model input, and the input feature sequence X is constructed. Each master candidate corresponds to a context feature sequence, and each context feature sequence is represented as a time-dependent multidimensional vector: ; Among them, x t is the feature at the tth moment, including a multi-dimensional context feature vector; ; in, The owner information in the metadata; Matching scores between fields and responsibilities; Timestamp for data upload; P peak is the peak value, tpeak is the corresponding time period; Faccess(u) is the frequency of data use; To modify preferences: Use the embedding layer to transform x t Encoded as a fixed-length vector e t , then the input feature sequence X is transformed into the embedded feature sequence E: ; The input sequence is the embedded feature sequence E, and the temporal context information is added through the positional encoding P to obtain the final input Z=E+P; Use the Transformer's self-attention mechanism to model the relationship between context features and calculate the weight of each time step : ; Where Q = ZW Q is the query matrix, K=ZW K is the bond matrix, V=ZW V is the value matrix; d k is the dimension of the query / key; G is a dense matrix; W Q、 W K、 W V are the weights of the query matrix, key matrix, and value matrix, respectively; is the activation function; Capturing dependencies between different features through a multi-head mechanism: ; in, is the weight coefficient; For the An attention head; represents the concatenation function; Output for the multi-head mechanism; ; in, For the attention head The trainable matrix of After l layers of Transformer encoding, the output feature sequence is obtained ; And use global pooling to aggregate sequence information to obtain a comprehensive recognition vector H = Pool (Z l ); The output is mapped to the category space of all candidate responsible persons through a fully connected layer: ; in, and are the weights and biases of the fully connected layer respectively; is the probability that the responsible person is u; Then, the ultimate person responsible for identification for: .

2. The method for optimizing power big data management in combination with data owners according to claim 1, characterized in that: The power big data includes power consumption behavior data, user operation behavior records and metadata; the power consumption behavior data includes user power consumption P(t), power grid equipment operation data, time series type data and special scenario records; the user operation behavior records are audit logs, including access, modification and reading frequency data; the metadata includes the table structure, data fields and field description information of the power internal management system.

3. The method for optimizing power big data management in combination with data owners according to claim 2 is characterized in that: The context features include power consumption behavior pattern values, operation behavior features and metadata characteristics, as follows: The power consumption behavior pattern value includes peak load period and abnormal fluctuation. The daily load value is calculated from the user's power consumption P(t) and the peak value P is extracted. peak and the corresponding period t peak : ; Calculate the difference between the user's electricity consumption at time t+1 and the user's electricity consumption at time t , if ΔP>T threshold , then the corresponding time t is marked as an abnormal fluctuation point, where T threshold To preset the threshold: ; Among them, mean and std are the fluctuation mean and standard deviation, and k is the adjustment coefficient; The operation behavior characteristics include data usage frequency F access (u) and modify preferences : ; ; Among them, N access (u) is the number of times user u accesses the data within time T; For user u for field f i The number of all modification operations, The total number of modifications made by user u on all fields; The metadata features include extracting field semantic relevance and naming patterns extracted from table names or field names; the field semantic relevance is calculated using cosine similarity: ; Among them, V( f 1 ),V( f 2 ) are fields f 1 ,f 2 The embedding vector of They are V( f 1 ),V( f 2 )'s mold length.

4. The method for optimizing power big data management in combination with data owners according to claim 1, characterized in that: The S2 is specifically: Use SQL to query the system's metadata table to collect metadata ownership and field description information; use the ELK log collection and processing framework to collect user operation logs in real time; Initialize the unified field structure of log data and metadata, associate table data by time and user ID, parse and classify field names, and classify field names into specific responsibility types; Prioritize matching the owner_id information in the metadata. If there is no owner record, match the metadata field description information or creator identity. If the metadata cannot be identified, identify the initial owner based on the upload log and the most recent upload operation.

5. The method for optimizing power big data management in combination with data owners according to claim 1, characterized in that: The dense matrix G is constructed based on GNN, as follows: The input feature sequence X is regarded as a node representation of the graph, and each time step feature x t Corresponding to a node; the initial adjacency matrix A is a fully connected graph, initialized to the unit matrix; Define the node update formula in the m-th layer GNN: ; in, is the hidden feature of node a in the mth layer; is the hidden feature of node b in the mth layer; Indicates the strength of the relationship between node a and node b; is the adjacent node of node a; , are the learnable weights and biases in GNN respectively; The node hidden features are updated in each layer, and the adjacency matrix is ​​adjusted according to the feature changes to obtain a dynamically updated adjacency matrix: ; in, is the activation function; is the similarity function; Finally, after the M-layer GNN update, the output adjacency matrix is ​​the dynamically generated dense matrix .

6. The method for optimizing power big data management in combination with data owners according to claim 1, characterized in that: The S4 is specifically: The data owner identification model processes the input data, generates ownership results, interacts with the data governance platform through interfaces, and transmits structured ownership data; The data governance platform completes data storage and permission allocation updates, and triggers the blockchain writing process; The blockchain records the identification process and ownership changes, and supports traceability and historical record query; The blockchain and the data governance platform database are synchronized to ensure that all ownership changes in the data governance platform are reflected on the chain, and the consistency of the platform data is verified from the historical records on the chain.

7. The method for optimizing power big data management in combination with data owners according to claim 1, characterized in that: The S5 is specifically: Build a rule base to store predefined detection rules and meta-information, configure anomaly detection models, and run them adaptively after parameterized training; Based on the rule base and anomaly detection model, data problems can be detected for scheduled or real-time scenarios; Generate a structured problem record for each detection problem, and organize the detected data problems into work orders that are readable by the data owner; The data owner is determined through the output of the data owner identification model, and the work order is pushed to the data owner to be responsible for repair or review.

8. The method for optimizing power big data management in combination with data owners according to claim 1, characterized in that: The S6 is specifically: The data owner receives the problem ticket through the ticket receiving platform, including the problem description, impact scope and repair suggestions; After the data owner repairs the data, he / she submits the repair results to the data governance platform; The data governance platform calls the compliance verification module to verify the repair results: If the repair is compliant, the work order is closed; if the repair is not compliant, a reverse work order is generated, and the problem and verification feedback are sent to the data owner for further processing.

Citation Information

Patent Citations

  • Regular expression-based acquisition, storage and analysis method of power big data

    CN104881424A

  • Power data management method and system based on data owner system

    CN119067593A