An AI-based power big data collection and processing method
By constructing a semantic propagation path graph of power maintenance logs and transaction contract fields, and identifying and correcting semantically offset field groups, the problem of field misalignment in power big data collection and processing was solved, thereby improving the accuracy of data reconciliation and the precision of scheduling archiving.
Patent Information
- Application Number
- CN202511439066.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2045-10-10
AI Technical Summary
Existing power big data collection and processing methods suffer from inconsistencies in content expression or mismatched context between fields when dealing with data from sources with significant differences, such as maintenance logs and transaction contracts. This leads to duplicate, missing, or mismatched archiving results, affecting the integrity of data reconciliation and the accuracy of scheduling archiving.
An AI-based approach is adopted to construct a semantic propagation path graph between power maintenance log fields and transaction contract fields using graph neural networks. This identifies and corrects semantically offset field groups, achieving synchronous correction of field semantic expression and sequential position, generating a semantically migrated path graph, and ensuring the accuracy of field connections and the continuity of time parameters.
It improves the semantic consistency and structural matching of power big data collection and processing results, solves the problems of field misalignment and reconciliation difficulties caused by heterogeneous sources of maintenance and transaction data, and improves the accuracy of data reconciliation and the precision of scheduling and archiving.
Smart Images

Figure CN120910037B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data acquisition technology, and in particular to an AI-based method for acquiring and processing large amounts of power data. Background Technology
[0002] The field of data acquisition technology involves the collection, transmission, and preliminary processing of various operation and maintenance data, transaction data, and related management information in power systems.
[0003] Among them, the power big data collection and processing method refers to the collection of raw data from specific devices such as electricity meters, load monitoring equipment, and transaction record terminals distributed in substations, distribution rooms, and user sides through a centralized collection system. Data is usually captured by periodic polling or event triggering at fixed time intervals and transmitted through serial communication or Ethernet. Subsequently, the data source is initially screened and structured by rule-based scripts in the main station or data center. Then, manually set field matching methods are used to organize and archive the electricity data, electricity price information, contract parameters, and other contents.
[0004] In the process of collecting and processing big data in the power industry, existing technologies mainly rely on rule scripts and field matching methods for preliminary screening and structuring. The processing methods lack the ability to dynamically adapt to semantic differences and structural misalignments between fields. When faced with data from sources with large differences, such as maintenance logs and transaction contracts, inconsistencies in content expression or mismatches in context often lead to duplicate, omission, or mismatch issues in the archived results. In particular, when there are non-standardized descriptions between contract terms and maintenance records, the semantics of fields cannot be effectively matched, which affects the completeness of data reconciliation and the accuracy of scheduling archives. If such archived data is relied upon before the execution of scheduling tasks, it is easy to cause equipment instruction deviations or contract execution risks. Summary of the Invention
[0005] The purpose of this invention is to address the shortcomings of existing technologies by proposing an AI-based method for collecting and processing large amounts of electricity data.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: an AI-based method for collecting and processing power big data, comprising the following steps:
[0007] S1: Collect power maintenance logs and power maintenance transaction contract data for power maintenance, extract power maintenance log fields and transaction contract fields from them, and perform sequential numbering on the upper and lower field positions of the maintenance log fields and transaction contract fields in their respective records to generate a set of basic field information.
[0008] S2: Using a graph neural network, semantic propagation paths are mapped to the maintenance log field and transaction contract field in the field basic information set to generate a semantic propagation path graph;
[0009] S3: Based on the semantic propagation path diagram, determine the differences in semantic expression between the maintenance log field and the transaction contract field, and filter the semantic offset field group;
[0010] S4: Based on the position of the maintenance log field and the transaction contract field above and below each other, perform bidirectional semantic propagation path correction on the semantic offset field group to generate a semantically migrated path diagram;
[0011] S5: Obtain the combination of all maintenance log fields and transaction contract fields that have completed semantic mapping and alignment of upper and lower field positions in the path graph after semantic migration, and obtain the power big data collection and processing results.
[0012] As a further embodiment of the present invention, the set of basic field information includes field content items, field position numbers, and field source identifiers; the semantic propagation path graph includes the sequential path of upper and lower fields, the semantic correspondence of fields, and time association information; the semantic offset field group includes structural offset field pairs, semantic mismatch field pairs, and field position difference values; the path graph after semantic migration includes path update results, field connection directions, and sequential position adjustment data; and the power big data collection and processing results specifically include field combination content and semantic mapping status.
[0013] As a further aspect of the present invention, the steps for obtaining the basic information set of the fields are specifically as follows:
[0014] S111: Collect power maintenance logs and power maintenance transaction contract data from power maintenance, extract maintenance log fields and transaction contract fields from them, and generate field extraction results;
[0015] S112: Based on the field extraction results, perform sequential numbering on the positions of the maintenance log field and the transaction contract field in their respective records, call the arrangement order of the fields in the original data to establish position index values, and bind the position index values with the corresponding field names to generate field order labeling data;
[0016] S113: Based on the field order, label the data, integrate all numbered maintenance log fields and transaction contract fields, and generate a set of basic field information.
[0017] As a further aspect of the present invention, the maintenance log fields include a description of the maintenance behavior, a maintenance device identifier, and a maintenance time.
[0018] As a further aspect of the present invention, the transaction contract fields include contract task description, contract equipment terms, and contract performance time.
[0019] As a further aspect of the present invention, the step of obtaining the semantic propagation path graph specifically includes:
[0020] S211: Using the set of basic information of the fields as the basis for connection in the graph neural network, construct and maintain the sequential connection path between the log field and the transaction contract field, set the edge connection relationship according to the order of the fields in the original data, and generate the field sequential connection structure.
[0021] S212: Based on the field sequential connection structure, extract the maintenance time and contract performance time as time edge association information, extract the maintenance equipment identifier and perform semantic correspondence judgment with the contract equipment clause, attach the semantic relationship of the matching fields to the corresponding connection path, and generate semantic structure association information;
[0022] S213: Construct a graph neural network graph structure based on the semantic structure association information, maintain log fields and transaction contract fields as nodes in the graph, connection paths as edges, semantic and time association information as edge attributes, and generate a semantic propagation path graph.
[0023] As a further aspect of the present invention, the step of obtaining the semantic offset field group specifically includes:
[0024] S311: Call the upper and lower field position index data of the maintenance log field and transaction contract field in the semantic propagation path diagram, extract the offset relationship based on the arrangement order of the upper and lower fields, identify the field combination with corresponding offset in the record, and generate upper and lower structure difference field pairs.
[0025] S312: Extract the semantic content of the maintenance equipment identifier and the contract equipment clause in the upper and lower structure difference field pair, perform key term matching and semantic consistency judgment, identify field combinations with no common referent in semantic expression, record semantic mismatch identifiers on the corresponding path, and generate semantic mismatch field pairs;
[0026] S313: Combining the above and below structural difference field pairs with the semantic mismatch field pairs, filter the field groups that simultaneously have semantic expression differences and structural offsets, as a combination of fields with inconsistent semantic expression and misaligned order, and generate a semantic offset field group.
[0027] As a further aspect of the present invention, the step of obtaining the semantically transferred path graph specifically includes:
[0028] S411: Call the upper and lower field positions of each maintenance log field and transaction contract field in the semantic offset field group, identify the field combination with upper and lower field position offset, and extract the time parameters corresponding to the field group as maintenance time and contract performance time to generate a structure corresponding field relationship group;
[0029] S412: Based on the field relationship group corresponding to the structure, determine the connection status of the maintenance equipment identifier and the contract equipment clause in the semantic propagation path graph of each group of fields, detect whether there is a path connection relationship, and determine whether the time is continuous based on the arrangement order of maintenance time and contract performance time, filter the field group that has been connected in the semantic propagation path graph and has time continuity, and generate a set of valid field groups for path connection.
[0030] S413: Adjust the connection direction and upper and lower field position numbers of the maintenance log field and transaction contract field in the semantic propagation path graph of the effective field set of the path connection respectively, update the connection order and path direction of the original fields in the graph, and generate the semantically migrated path graph.
[0031] As a further aspect of the present invention, the steps for obtaining the power big data collection and processing results are specifically as follows:
[0032] S511: Call the connection data of the maintenance log field and transaction contract field in the path graph after semantic migration, filter all field combinations that have completed semantic mapping and whose upper and lower field positions have been aligned, and generate a list of aligned field combinations;
[0033] S512: Based on the aligned field combination list, organize the content, upper and lower field position numbers and semantic connection direction of each group of fields, map and archive the combination relationship between the maintenance behavior description and the contract task description, and the maintenance equipment identifier and the contract equipment clause in the structure diagram, and generate field comparison structure data.
[0034] S513: Based on the field comparison structure data, the field content and structural relationship are used as input fields for power maintenance and transaction data reconciliation and scheduling archiving, generating power big data collection and processing results.
[0035] Compared with the prior art, the advantages and positive effects of the present invention are as follows:
[0036] In this invention, by extracting the field content and their relative positions from power maintenance logs and transaction contract data using numbering, a propagation path between fields is constructed at the semantic level. After identifying differences, field groups with semantic offsets and structural misalignments are bidirectionally aligned, enabling synchronous correction of field semantic expression and sequential position during the association process. This improves the accuracy of field connections between different types of data. Combined with time parameters for continuity judgment, the connection between maintenance and contract fields is made more aligned with actual business processes. After aligning the field combinations, the field content and connection relationships are further archived, providing a structured basis for data reconciliation and scheduling aggregation. This processing flow significantly improves the semantic consistency, structural matching degree, and application usability of power big data collection and processing results, solving problems such as field misalignment and reconciliation difficulties caused by heterogeneous sources of maintenance and transaction data. Attached Figure Description
[0037] Figure 1 This is a schematic diagram of the main steps of the present invention;
[0038] Figure 2 This is a flowchart of step S1 of the present invention;
[0039] Figure 3 This is a flowchart of step S2 of the present invention;
[0040] Figure 4 This is a flowchart of step S3 of the present invention;
[0041] Figure 5 This is a flowchart of step S4 of the present invention;
[0042] Figure 6 This is a flowchart of step S5 of the present invention. Detailed Implementation
[0043] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0044] Please see Figure 1 This invention provides a technical solution: an AI-based method for collecting and processing big data on electricity, comprising the following steps:
[0045] S1: Collect power maintenance logs and power maintenance transaction contract data for power maintenance, extract power maintenance log fields and transaction contract fields from them, and perform sequential numbering on the upper and lower field positions of the maintenance log fields and transaction contract fields in their respective records to generate a set of basic field information.
[0046] S2: Using a graph neural network, semantic propagation paths are mapped to the maintenance log field and transaction contract field in the field basic information set to generate a semantic propagation path graph;
[0047] S3: Based on the semantic propagation path graph, determine the differences in semantic expression between maintenance log fields and transaction contract fields, and filter semantic offset field groups;
[0048] S4: Based on the position of the maintenance log field and the transaction contract field above and below each other, perform bidirectional semantic propagation path correction on the semantic offset field group to generate a path diagram after semantic migration;
[0049] S5: Obtain the combination of all maintenance log fields and transaction contract fields that have completed semantic mapping and alignment of upper and lower field positions in the path graph after semantic migration, and obtain the power big data collection and processing results;
[0050] The set of basic field information includes field content items, field location numbers, and field source identifiers. The semantic propagation path diagram includes the sequential path of upper and lower fields, the semantic correspondence of fields, and time association information. The semantic offset field group includes structural offset field pairs, semantic mismatch field pairs, and field position difference values. The path diagram after semantic migration includes path update results, field connection directions, and sequential position adjustment data. The specific results of power big data collection and processing are field combination content and semantic mapping status.
[0051] Please see Figure 2 The specific steps for obtaining the basic information set of fields are as follows:
[0052] S111: Collect power maintenance logs and power maintenance transaction contract data from power maintenance, extract maintenance log fields and transaction contract fields from them, and generate field extraction results;
[0053] For power maintenance logs collected from the power maintenance server and power maintenance transaction contract data collected from the contract management database, the unstructured text of the power maintenance logs is scanned and matched line by line using a pre-defined keyword list containing "maintenance time," "maintenance personnel," "maintenance location," "maintenance equipment," "fault description," and "handling process." Once a keyword in the list is matched, all characters from the beginning of the keyword to the end of the line are extracted as the field content for that keyword. For the structured documents of the power maintenance transaction contracts, "contract number" and "signing date" are extracted directly from specific locations, such as the "basic contract information" section, and "contract equipment terms" and "performance time" are extracted from the "contract terms" section, based on a predefined document template. All field names and their contents extracted from the power maintenance logs and power maintenance transaction contract data are then aggregated to generate the field extraction results.
[0054] S112: Based on the field extraction results, perform sequential numbering on the positions of the maintenance log field and the transaction contract field in their respective records, call the arrangement order of the fields in the original data to establish position index values, and bind the position index values with the corresponding field names to generate field order labeling data;
[0055] For the field extraction results, the maintenance log field, which originates from a single power maintenance log record, is assigned a position index value starting from 1 and continuously increasing, based on the order in which each field appears in the original log text from top to bottom. Each position index value is then bound to the corresponding field name, forming a pair of data containing the index and the name. Similarly, the transaction contract field, which originates from a single power maintenance transaction contract, is assigned a position index value starting from 1 and continuously increasing, based on the order in which each field is arranged in the contract document structure. Each position index value is then bound to the corresponding field name. The above numbering and binding operations are performed on all maintenance log fields and transaction contract fields to generate field order labeling data.
[0056] S113: Label the data according to the field order, integrate all numbered maintenance log fields and transaction contract fields, and generate a set of basic field information;
[0057] For data with sequentially labeled fields, an integration process is initiated. This process iterates through all maintenance log fields and transaction contract fields labeled with location index values. Each field is treated as an independent data entry, along with its source type (i.e., maintenance log or transaction contract), location index value, and name. This entire process is then moved into a newly created set. The integration process does not modify any existing field attributes; it only performs aggregation actions until all labeled fields in all records are included in the newly created set, forming a set of basic field information that includes all numbered maintenance log fields and transaction contract fields.
[0058] Please see Figure 3 The specific steps for obtaining the semantic propagation path graph are as follows:
[0059] S211: Using the set of basic field information as the basis for connections in the graph neural network, construct and maintain the sequential connection path between log fields and transaction contract fields, set the edge connection relationship according to the order of the fields in the original data, and generate the field sequential connection structure.
[0060] Based on the set of basic field information, each field in the set is treated as an independent node. Between nodes originating from the same maintenance log record, a directed connection edge is established from the node with the smaller value to the next larger value based on the node's position index value. The same operation is performed between nodes originating from the same transaction contract, that is, internal directed connection paths are established according to the node's position index value in ascending order. During the connection establishment process, no connections are created that cross different maintenance log records or different transaction contracts, nor are connections established between maintenance log field nodes and transaction contract field nodes, thus generating a field sequential connection structure.
[0061] S212: Based on the field sequential connection structure, extract the maintenance time and contract performance time as time edge association information, extract the maintenance equipment identifier and perform semantic correspondence judgment with the contract equipment clause, and attach the semantic relationship of the matching fields to the corresponding connection path to generate semantic structure association information;
[0062] In the field sequential join structure, first, all nodes named "Maintenance Time" and all nodes named "Contract Performance Time" are searched, and the time information contained in these two nodes is extracted and marked as time edge association information. Then, the content of all "Maintenance Equipment Identifier" nodes and all "Contract Equipment Terms" nodes are extracted. Semantic correspondence judgment is performed on the content strings of these two types of nodes. The judgment is completed by calculating the Jaccard similarity coefficient, the formula of which is: The specific meaning of each letter in the formula is as follows: : Represents the Jaccard similarity coefficient between the content of the "Maintenance Equipment Identifier" field and the "Contract Equipment Terms" field. It is a final result that quantifies the degree of similarity between two text strings. : This represents a set of characters formed after deduplication of the "Maintenance Equipment Identifier" field content from the power maintenance log. Each element in the set is a unique character from the field content. : This represents a set of characters formed after deduplication of the "Contract Equipment Terms" field content from the power maintenance transaction contract. Each element in the set is a unique character from the field content. : Represents a set With sets The number of elements in the intersection, sign This represents the intersection operation, which filters out elements that exist in both sets. and set All common characters, symbols This indicates the total number of characters contained in the intersection. : Represents a set With sets The number of elements in a union, sign This represents the union operation, which merges sets. and set All characters and remove duplicates, symbols This represents the total number of characters contained in the union of sets. For example, if the content of "Maintenance Equipment Identifier" is "#3 Main Transformer", then its character set... Given {'#','3','main','transformer','voltage','device'}, if the content of the "Contract Equipment Terms" is "Maintenance of #3 Main Transformer", then its character set is... For {'#','3','main','transformer','voltage','equipment','maintenance','protection'}, first calculate the intersection. ={'#','3','main','transformer','voltage','device'}, its number of elements The value is 6, then the union is calculated. ={'#','3','Main','Transformer','Voltage','Electrical','Maintenance','Protection'}, its number of elements The value is 8. Substitute the value into the formula: Set a similarity threshold for semantic correspondence judgment. The threshold setting process is as follows: Prepare a verification dataset containing several pairs of device description fields. Power experts manually label each pair of fields as "corresponding" or "not corresponding". Then, calculate the Jaccard similarity coefficient for each pair of fields in the dataset. For all possible thresholds from 0.1 to 0.9, calculate the true positive rate and false positive rate for each threshold. Finally, select the coefficient value that maximizes the difference between the true positive rate and the false positive rate as the final threshold. Since the calculated similarity coefficient of 0.75 is greater than the threshold of 0.7, it is determined that the two fields are semantically matched. The matching conclusion is attached as a semantic relationship to the possible connection path to be established in the future to generate semantic structure association information.
[0063] S213: Construct a graph neural network graph structure based on semantic structure association information, maintain log fields and transaction contract fields as nodes in the graph, connection paths as edges, semantic and time association information as edge attributes, and generate a semantic propagation path graph;
[0064] By leveraging semantic structure association information, we first treat all maintenance log fields and transaction contract fields in the basic field information set as nodes in the graph structure, and the sequential connection paths in the field sequential connection structure as the basic edges between nodes. Then, for each pair of "maintenance equipment identifier" nodes and "contract equipment terms" nodes that are marked as semantically matched in the semantic structure association information, we add an edge between the two nodes to indicate semantic association, and assign the attribute "semantic relationship: match" to the newly added edge. At the same time, we find the "maintenance time" node and "contract performance time" node that are in the same original record as these two matching nodes, add an edge between these two time nodes, and assign the extracted time information as the "time association" attribute to this time edge, thus generating a semantic propagation path graph.
[0065] Please see Figure 4 The specific steps for obtaining the semantic offset field group are as follows:
[0066] S311: Call the index data of the upper and lower field positions of the log field and the transaction contract field in the semantic propagation path graph, extract the offset relationship based on the arrangement order of the upper and lower fields, identify the field combination with the corresponding offset in the record, and generate the upper and lower structure difference field pair.
[0067] Based on the semantic propagation path graph and field order annotation data, we first filter all maintenance log field node and transaction contract field node pairs that are directly connected by edges with the "semantic relationship: matching" attribute in the semantic propagation path graph. For each selected pair of nodes, we query the corresponding upper and lower field position index values from the field order annotation data. Then, we compare the two position index values. If the two position index values are not equal, the field pair is identified as a combination with corresponding position offsets. All the identified field combinations with corresponding position offsets are then collected to generate upper and lower structure difference field pairs.
[0068] S312: Extract the semantic content of maintenance equipment identifier and contract equipment clause in the upper and lower structure difference field pairs, perform key term matching and semantic consistency judgment, identify field combinations with no common referent in semantic expression, record semantic mismatch identifiers on the corresponding paths, and generate semantic mismatch field pairs;
[0069] For pairs of fields with structural differences, the original semantic content of each pair of "Maintenance Equipment Identifier" and "Contract Equipment Terms" fields is extracted. Using a word vector model pre-trained on a massive power industry document corpus, key terms in the field content are quantified into numerical vectors. For fields containing multiple key terms, the element-wise average of all key term vectors is taken as the semantic vector for the entire field. Then, the semantic vectors of the two fields are calculated... and The cosine distance between them is used to determine whether there is a semantic mismatch. The calculation formula is as follows: The specific meaning of each letter in the formula is as follows: : Represents a vector and The cosine distance between two fields is the final calculated result used to measure the degree of semantic dissimilarity between them. : Represents the semantic vector obtained after the "Maintenance Equipment Identifier" field is transformed by the word vector model. This vector is a set of values that captures the specific meaning of the field in the power industry. : This represents the semantic vector obtained after the "Contract Equipment Terms" field has been transformed using the same word vector model. This vector is also a set of numerical values used to mathematically express the semantics of the field. : Represents semantic vector In the Each dimension contains specific numerical components, and each dimension corresponds to an abstract semantic feature learned by the model. : Represents semantic vector In the The specific numerical components in each dimension, and They correspond to the same semantic feature dimension. : The index number representing the dimension of the vector, starting from 1 and continuing up to the total number of dimensions of the vector. This is used to iterate through each component of the vector. : Represents the total number of dimensions of the vector space defined by the word vector model. It is a fixed integer, such as 300, which determines the number of numerical components contained in each semantic vector. : Represents the summation operation, indicating that the expression immediately following it will be summed from... arrive The summation of all calculation results. Taking an example, if the average vector of the field "Replace #3 transformer bushing" is... for The average vector of the field "Inspecting C-line switch" for First, calculate the dot product: ,calculate Norm: ,calculate Norm: Calculate the cosine similarity: Finally, calculate the cosine distance: Set a distance threshold for semantic mismatch. The threshold is set to 0.6. The basis for setting the threshold is to perform distance calculation on a validation set containing a large number of power term pairs to form distance distributions of two types of term pairs: "relevant" and "unrelevant". The distance value at the intersection of the two distribution curves is selected as the threshold because the point can minimize the misjudgment of the two types of terms. Since the calculated distance of 0.6087 is greater than the threshold of 0.6, it is determined that the semantic expressions of these two fields have no common referents. A semantic mismatch identifier is added to the path connecting these two fields in the semantic propagation path graph. All the fields that have gone through the above judgment process and have been recorded with semantic mismatch identifiers are combined to generate semantic mismatch field pairs.
[0070] S313: Combining the structural difference field pairs with the semantic mismatch field pairs, filter the field groups that have both semantic expression differences and structural offsets, and generate semantic offset field groups as a combination of fields with inconsistent semantic expression and misaligned order.
[0071] Based on the structural difference field pairs and semantic mismatch field pairs, a filtering process is initiated. Each field combination in the structural difference field pairs is traversed, and it is checked whether the currently traversed field combination also exists in the set of semantic mismatch field pairs. If a field combination satisfies both the conditions of existing in the structural difference field pairs and existing in the semantic mismatch field pairs, that is, the field combination has both differences in position index values and semantic mismatch, then the field combination is filtered out. All the filtered field combinations are used to generate a semantic offset field group.
[0072] Please see Figure 5 The specific steps for obtaining the semantically transferred path graph are as follows:
[0073] S411: Call the upper and lower field positions of each maintenance log field and transaction contract field in the semantic offset field group, identify the field combinations with upper and lower field position offsets, and extract the time parameters corresponding to the field group as maintenance time and contract performance time to generate the structure corresponding field relationship group;
[0074] For the semantic offset field group, the internal structure of each maintenance log field and transaction contract field is analyzed to identify and extract subordinate fields with the same name. It is then determined whether these subordinate fields are offset in relative position within their respective parent fields. If an offset is identified, position alignment is performed according to their order in the parent field content to ensure consistent relative positions. Subsequently, if both the maintenance log field and transaction contract field groups contain subordinate field combinations with the same name and aligned relative positions, the time parameters of these subordinate fields are extracted and defined as maintenance time and contract fulfillment time, respectively. All field combinations that have achieved structural alignment and successfully extracted time parameters through the above method are combined to generate a structurally corresponding field relationship group.
[0075] S412: Based on the corresponding field relationship group of the structure, determine the connection status of the maintenance equipment identifier and the contract equipment clause in the semantic propagation path graph of each group of fields, detect whether there is a path connection relationship, and determine whether the time is continuous based on the arrangement order of maintenance time and contract performance time. Filter the field groups that have been connected in the semantic propagation path graph and have time continuity, and generate a set of valid field groups for path connection.
[0076] For each pair of maintenance equipment identifier and contract equipment clause fields in the structure corresponding to the field relationship group, the connection status is determined in the semantic propagation path graph. The determination process is to perform a path search in the graph starting from a field node and check whether there is a path consisting of edges that can reach another field node. If such a path exists, the path connection relationship is determined to exist. Subsequently, the time continuity of the maintenance time and contract performance time corresponding to the field group is determined. The determination process is to convert the two time values into a comparable numerical format and then compare their size. If the value of the maintenance time is not less than the value of the contract performance time, the time is determined to be continuous. Only those field groups that have established connections in the semantic propagation path graph and have time continuity are selected to generate a valid field set for path connection.
[0077] S413: Adjust the connection direction and the position number of the upper and lower fields in the semantic propagation path graph for the maintenance log field and the transaction contract field in the valid field set of the path connection, respectively, update the connection order and path direction of the original fields in the graph, and generate the semantically migrated path graph.
[0078] For each field group in the valid field set of path connections, an adjustment operation is performed in the data structure of the semantic propagation path graph. First, the edge between the maintenance log field node and the transaction contract field node in the connection field group is located, and the direction attribute of the edge is accessed. If the direction attribute is from the maintenance log field to the transaction contract field, or is undirected, the attribute is modified to be from the transaction contract field to the maintenance log field. This adjustment establishes the logical dominant relationship from the power maintenance transaction contract to the power maintenance log. Next, the position index value attribute of the transaction contract field node in the field group is obtained, and then the maintenance log field node is located. Its position index value attribute is updated to the position index value of the transaction contract field that was just obtained. This operation makes the two nodes completely aligned in structural order, avoiding misalignment caused by different data sources. The above edge direction modification and node position index value update operations are repeated for all field groups in the valid field set of path connections, thereby updating the connection order and path direction of the original fields in the graph and generating a semantically migrated path graph.
[0079] Please see Figure 6 The specific steps for obtaining the results of power big data collection and processing are as follows:
[0080] S511: Call the connection data between the maintenance log field and the transaction contract field in the path graph after semantic migration, filter all field combinations that have completed semantic mapping and whose upper and lower field positions have been aligned, and generate a list of aligned field combinations;
[0081] In the semantically transferred path graph, a filtering program is initiated. The program traverses each edge in the graph and performs a series of checks on each edge. First, it checks the two nodes connected by the edge. One node's type attribute must be "maintenance log field," and the other node's type attribute must be "transaction contract field." If this condition is not met, the current edge is skipped. Second, it checks the edge's attributes. It must contain a "semantic relationship" attribute with a value of "match" to confirm that the two fields have been semantically validated and confirmed to be related. If this condition is not met, it is also skipped. Finally, it retrieves the "position index value" attribute of the two nodes connected by the edge and compares whether the two attribute values are completely equal to confirm that the two fields are aligned in structural order. If they are not equal, it is also skipped. Only when an edge passes all three checks will the program extract the maintenance log field and transaction contract field connected by this edge as a valid combination. All extracted valid combinations are added to a new list, ultimately generating a list of aligned field combinations.
[0082] S512: Based on the aligned field combination list, organize the content, upper and lower field position numbers and semantic connection direction of each group of fields, map and archive the combination relationship between the maintenance behavior description and the contract task description, and the maintenance equipment identifier and the contract equipment clause in the structure diagram, and generate field comparison structure data.
[0083] For each group of fields in the aligned field combination list, a structured organization and archiving process is executed. The process creates a new data record for each group of fields. In the new record, the content of the power maintenance log field node is extracted and stored in the "Maintenance Behavior Description" field, and the content of the power maintenance transaction contract field node is extracted and stored in the "Contract Task Description" field. Then, the aligned "Location Index Value" is obtained from any node and stored in the "Structural Location Number" field. The direction attribute of the edge connecting the two nodes is obtained and stored in the "Semantic Connection Direction" field. In this way, the combination relationship between the maintenance behavior description and the contract task description, and between the maintenance equipment identifier and the contract equipment clause in the structure diagram is mapped one by one. These structured data records containing complete comparison information are archived to generate field comparison structure data.
[0084] S513: Based on field-to-structure data, the content and structure relationship of fields are used as input fields for power maintenance and transaction data reconciliation and scheduling archiving, generating power big data collection and processing results;
[0085] Using the field-matching structured data as input, the system is used for power maintenance and transaction data reconciliation and scheduling archiving. In the reconciliation process, the automated program reads each record in the field-matching structured data and directly compares the content of the "Maintenance Behavior Description" and "Contract Task Description" fields. If the content is determined to be consistent under preset rules, such as identical characters or matching key technical parameters, the record is marked as "Reconciliation Successful"; otherwise, it is marked as "Pending Manual Review". In the scheduling archiving process, the program uses the "Structure Location Number" in the records to uniformly sort the fields originating from power maintenance logs and power maintenance transaction contracts, and establishes logical links between archived data based on the "Semantic Connection Direction". This integrates the originally scattered maintenance log data and contract data into a consistent, logically clear, and traceable archive set. The outputs of these two processes together constitute the power big data collection and processing results.
[0086] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention in any other way. Any person skilled in the art may make changes or modifications to the above-disclosed technical content to create equivalent embodiments that can be applied to other fields. However, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A method for collecting and processing big data on electricity based on AI, characterized in that, Includes the following steps: S1: Collect power maintenance logs and power maintenance transaction contract data for power maintenance, extract power maintenance log fields and transaction contract fields from them, and perform sequential numbering on the upper and lower field positions of the maintenance log fields and transaction contract fields in their respective records to generate a set of basic field information. S2: Using a graph neural network, semantic propagation paths are mapped to the maintenance log field and transaction contract field in the field basic information set to generate a semantic propagation path graph; S3: Based on the semantic propagation path diagram, determine the differences in semantic expression between the maintenance log field and the transaction contract field, and filter the semantic offset field group; S4: Based on the position of the maintenance log field and the transaction contract field above and below each other, perform bidirectional semantic propagation path correction on the semantic offset field group to generate a semantically migrated path diagram; S5: Obtain the combination of all maintenance log fields and transaction contract fields that have completed semantic mapping and alignment of upper and lower field positions in the path graph after semantic migration, and obtain the power big data collection and processing results; The specific steps for obtaining the basic information set of the fields are as follows: S111: Collect power maintenance logs and power maintenance transaction contract data from power maintenance, extract maintenance log fields and transaction contract fields from them, and generate field extraction results; S112: Based on the field extraction results, perform sequential numbering on the positions of the maintenance log field and the transaction contract field in their respective records, call the arrangement order of the fields in the original data to establish position index values, and bind the position index values with the corresponding field names to generate field order labeling data; S113: Based on the field order labeling data, integrate all numbered maintenance log fields and transaction contract fields to generate a set of basic field information; The steps for obtaining the semantic propagation path graph are as follows: S211: Using the set of basic information of the fields as the basis for connection in the graph neural network, construct and maintain the sequential connection path between the log field and the transaction contract field, set the edge connection relationship according to the order of the fields in the original data, and generate the field sequential connection structure. S212: Based on the field sequential connection structure, extract the maintenance time and contract performance time as time edge association information, extract the maintenance equipment identifier and perform semantic correspondence judgment with the contract equipment clause, attach the semantic relationship of the matching fields to the corresponding connection path, and generate semantic structure association information; S213: Construct a graph neural network graph structure based on the semantic structure association information, maintain log fields and transaction contract fields as nodes in the graph, connection paths as edges, semantic and time association information as edge attributes, and generate a semantic propagation path graph; The specific steps for obtaining the semantic offset field group are as follows: S311: Call the upper and lower field position index data of the maintenance log field and transaction contract field in the semantic propagation path diagram, extract the offset relationship based on the arrangement order of the upper and lower fields, identify the field combination with corresponding offset in the record, and generate upper and lower structure difference field pairs. S312: Extract the semantic content of the maintenance equipment identifier and the contract equipment clause in the upper and lower structure difference field pair, perform key term matching and semantic consistency judgment, identify field combinations with no common referent in semantic expression, record semantic mismatch identifiers on the corresponding path, and generate semantic mismatch field pairs; S313: Combining the above and below structural difference field pairs with the semantic mismatch field pairs, filter the field groups that simultaneously have semantic expression differences and structural offsets, as a combination of fields with inconsistent semantic expression and misaligned order, and generate a semantic offset field group.
2. The AI-based power big data acquisition and processing method according to claim 1, characterized in that, The set of basic field information includes field content items, field position numbers, and field source identifiers. The semantic propagation path graph includes the sequential paths of upper and lower fields, the semantic correspondence of fields, and time association information. The semantic offset field group includes structural offset field pairs, semantic mismatch field pairs, and field position difference values. The path graph after semantic migration includes path update results, field connection directions, and sequential position adjustment data. The power big data collection and processing results specifically include field combination content and semantic mapping status.
3. The AI-based power big data acquisition and processing method according to claim 1, characterized in that, The maintenance log fields include a description of the maintenance activity, the identifier of the device being maintained, and the maintenance time.
4. The AI-based power big data acquisition and processing method according to claim 1, characterized in that, The transaction contract fields include contract task description, contract equipment terms, and contract performance time.
5. The AI-based power big data acquisition and processing method according to claim 1, characterized in that, The specific steps for obtaining the semantically transferred path graph are as follows: S411: Call the upper and lower field positions of each maintenance log field and transaction contract field in the semantic offset field group, identify the field combination with upper and lower field position offset, and extract the time parameters corresponding to the field group as maintenance time and contract performance time to generate a structure corresponding field relationship group; S412: Based on the field relationship group corresponding to the structure, determine the connection status of the maintenance equipment identifier and the contract equipment clause in the semantic propagation path graph of each group of fields, detect whether there is a path connection relationship, and determine whether the time is continuous based on the arrangement order of maintenance time and contract performance time, filter the field group that has been connected in the semantic propagation path graph and has time continuity, and generate a set of valid field groups for path connection. S413: Adjust the connection direction and upper and lower field position numbers of the maintenance log field and transaction contract field in the semantic propagation path graph of the effective field set of the path connection respectively, update the connection order and path direction of the original fields in the graph, and generate the semantically migrated path graph.
6. The AI-based power big data acquisition and processing method according to claim 5, characterized in that, The specific steps for obtaining the power big data collection and processing results are as follows: S511: Call the connection data of the maintenance log field and transaction contract field in the path graph after semantic migration, filter all field combinations that have completed semantic mapping and whose upper and lower field positions have been aligned, and generate a list of aligned field combinations; S512: Based on the aligned field combination list, organize the content, upper and lower field position numbers and semantic connection direction of each group of fields, map and archive the combination relationship between the maintenance behavior description and the contract task description, and the maintenance equipment identifier and the contract equipment clause in the structure diagram, and generate field comparison structure data. S513: Based on the field comparison structure data, the field content and structural relationship are used as input fields for power maintenance and transaction data reconciliation and scheduling archiving, generating power big data collection and processing results.
Citation Information
Patent Citations
Financial index real-time analysis method and system
CN120107003A