Power transmission and transformation project-oriented GIM model multi-format automatic conversion method
By identifying and removing invalid data in GIM model data, and using dynamic mapping rules and machine learning algorithms to optimize data mapping, the problems of low conversion efficiency and format incompatibility in existing GIM model data technologies are solved, achieving efficient and accurate data conversion.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-16
- Publication Date
- 2026-04-14
AI Technical Summary
Existing data conversion methods for GIM models are inefficient and susceptible to human error when dealing with complex power transmission and transformation engineering facilities, leading to data loss or format incompatibility.
By collecting GIM model data, identifying data formats and labeling types, eliminating invalid data, optimizing data mapping using dynamic mapping rules and machine learning algorithms, performing format compatibility checks and error repairs, and generating a conversion report.
It improves the accuracy and consistency of data transformation, enhances the efficiency and accuracy of data processing, and provides an efficient and flexible automated data transformation solution.
Smart Images

Figure CN121858645A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of automated conversion technology, and in particular to an automatic multi-format conversion method for GIM models in power transmission and transformation engineering. Background Technology
[0002] With the continuous development of information technology and engineering management needs, especially in the field of power transmission and transformation engineering, the complexity of data processing and transformation is gradually increasing, and industrial data processing has become a key factor. As an important three-dimensional geometric and spatial data representation method, GIM model is widely used in power transmission and transformation engineering to represent and manage various facilities, structures and their attribute data; the application of GIM model is gradually expanding to collaborative work between various engineering facilities and different data formats.
[0003] While existing technologies have made some progress in the construction and conversion of GIM models, they generally have some limitations. Traditional GIM model data conversion methods often rely on manual processes for data format conversion and mapping. When dealing with complex power transmission and transformation engineering facilities, these methods are inefficient and susceptible to human error, leading to data loss or format incompatibility. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides an automatic multi-format conversion method for GIM models in power transmission and transformation projects, which solves the problems of low efficiency and format incompatibility in the semi-automatic conversion process.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: This invention provides an automatic multi-format conversion method for GIM models in power transmission and transformation engineering, which includes: Collect GIM model data, identify the format of the source data in the GIM model data and mark the data type, and generate the original dataset; Remove invalid data from the original dataset and convert it into a unified standard format to generate a standardized dataset; Dynamic mapping rules are used to map power transmission and transformation engineering facilities and structural components in the standardized dataset to corresponding entities in the target format, generating a mapping dataset; By using machine learning algorithms, the mapping relationships and data structures of the mapping dataset are adjusted to obtain an optimized mapping dataset. The optimized mapping dataset is then subjected to format compatibility checks and error correction to generate a validation dataset. Based on the validation dataset, the differences between the source data and the target data in terms of geometry, attribute data, and topology are compared. Inconsistencies are detected and corrected step by step using a format compatibility verification tool, and a conversion report is generated.
[0007] As a preferred embodiment of the automatic multi-format conversion method for GIM models in power transmission and transformation engineering described in this invention, the specific steps for collecting GIM model data, identifying the format of the source data in the GIM model data, and marking the data type are as follows. Parse the data metadata of GIM model data, identify and verify data format consistency, and generate format identification information; Based on the format recognition information, analyze the source, physical meaning and data fields of each data element, obtain the data type label, and perform format conversion on the data type label to generate an intermediate dataset; Perform integrity checks on the intermediate dataset and generate integrity verification information.
[0008] As a preferred embodiment of the automatic multi-format conversion method for GIM models for power transmission and transformation engineering described in this invention, the generation of the original dataset refers to checking and repairing missing entries in the integrity verification information.
[0009] As a preferred embodiment of the automatic multi-format conversion method for GIM models in power transmission and transformation engineering described in this invention, the specific steps for removing invalid data from the original dataset and converting it into a unified standard format to generate a standardized dataset are as follows. Based on the original dataset, check whether the core fields of each data element are complete, whether the data values are within the valid range, and whether the format meets the standards, and remove invalid data to generate a clean dataset; Perform format conversion on the cleaned dataset and standardize the intermediate dataset; Perform integrity verification on the standardized intermediate dataset, repair missing data, and generate a standardized dataset.
[0010] As a preferred embodiment of the automatic multi-format conversion method for GIM models in power transmission and transformation engineering described in this invention, the method employs dynamic mapping rules to map power transmission and transformation engineering facilities and structural components in the standardized dataset to corresponding entities in the target format. The specific steps are as follows: By employing dynamic mapping rules, each element in the standardized dataset is matched with the corresponding entity in the target format, and a mapping relationship is established to generate preliminary mapping rules. Based on the preliminary mapping rules, analyze the specific conversion relationship between each data field in the standardized dataset and the entity attributes in the target format, and generate data mapping relationships; Based on the data mapping relationship, the matching rules for field names and attribute types are dynamically adjusted to generate the adjusted mapping rules; Based on the adjusted mapping rules, each element in the standardized dataset is mapped to the corresponding entity in the target format to generate a preliminary mapping dataset; Based on the initial mapping dataset, the weights of each data field during the mapping process are adjusted, and the mapping weight adjustment data is output.
[0011] As a preferred embodiment of the automatic multi-format conversion method for GIM models for power transmission and transformation engineering described in this invention, the generation of the mapping dataset refers to adjusting the data according to the mapping weight, mapping each entity in the standardized dataset according to the target format requirements, and performing type conversion and unit unification.
[0012] As a preferred embodiment of the automatic multi-format conversion method for GIM models in power transmission and transformation engineering described in this invention, the step of adjusting the mapping relationship and data structure of the mapping dataset using machine learning algorithms to obtain an optimized mapping dataset includes the following specific steps. The decision tree algorithm is used to extract features, normalize and encode one-hot data in the mapping dataset, and convert the numerical and categorical fields in the mapping dataset into a unified format to generate a converted dataset. Adjust the feature weights and data structure in the transformed dataset to obtain the adjusted transformed dataset, and predict the mapping relationship between data fields in the adjusted transformed dataset to generate the optimized mapped dataset.
[0013] As a preferred embodiment of the automatic multi-format conversion method for GIM models in power transmission and transformation engineering described in this invention, the specific steps for performing format compatibility checks and error corrections on the optimized mapping dataset to generate a verification dataset are as follows. The table structure in the optimized mapping dataset is transformed into a graph structure, the updated data structure is obtained, and the format compatibility of the updated data structure is checked to generate format verification data. Perform field data repair and unit conversion on the format validation data to generate a validation dataset.
[0014] As a preferred embodiment of the automatic multi-format conversion method for GIM models in power transmission and transformation engineering described in this invention, the following steps are taken: Based on the verification dataset, the differences between the source data and the target data in terms of geometry, attribute data, and topology are compared. Inconsistencies are then gradually detected and corrected using a format compatibility verification tool. Geometric information, attribute data, and topological structure of source and target data are extracted from the validation dataset, and a geometric matching algorithm is used to compare the differences one by one to generate a difference comparison report. Based on the difference comparison report, analyze the format issues between the source data and the target data, and perform field type, unit and precision checks to generate format compatibility verification data; Based on the format compatibility verification data, geometric anomalies are detected, and the data position and size are adjusted to generate a geometric repair dataset; Based on the geometric repair dataset, the mismatch of attribute data is handled, and the data format and units are unified to generate the attribute repair dataset; Based on the attribute repair dataset, the defects in the topology structure are resolved, the device connection relationships are adjusted, the missing topology connections are repaired, and a topology repair dataset is generated.
[0015] As a preferred embodiment of the automatic multi-format conversion method for GIM models for power transmission and transformation engineering described in this invention, the generation of the conversion report refers to recording the repair process of geometry, attribute data and topology based on the topology repair dataset.
[0016] The beneficial effects of this invention are as follows: by collecting and identifying GIM model data formats, invalid data is removed and converted into a standard format; dynamic mapping rules and machine learning are used to optimize data mapping; format compatibility checks and error repairs are performed; industrial data processing plays a key role in this process, ensuring the accuracy and consistency of data conversion; by comparing and repairing the differences between source data and target data in geometry, attributes, and topology, the efficiency and accuracy of data processing are greatly improved, providing an efficient and flexible automated data conversion solution. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart of a method for automatic multi-format conversion of GIM models for power transmission and transformation projects.
[0019] Figure 2 A flowchart for data identification and intermediate data generation.
[0020] Figure 3 The flowchart for executing dynamic mapping rules.
[0021] Figure 4 This is a flowchart for multidimensional difference comparison and repair. Detailed Implementation
[0022] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0023] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0024] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0025] Reference Figures 1-4 This is one embodiment of the present invention, which provides an automatic multi-format conversion method for GIM models in power transmission and transformation engineering, including the following steps: S1. Collect GIM model data, identify the format of the source data in the GIM model data and mark the data type, and generate the original dataset.
[0026] S1.1. Parse the data element information of the GIM model data, identify and verify the consistency of the data format, and generate format identification information.
[0027] It should be noted that each data element in the GIM model is extracted to determine its structure and format. Information such as type, field name, and unit is extracted from the GIM model data to confirm its data type (e.g., integers and floating values) and corresponding format requirements. Each data element is verified to conform to the format standards, ensuring that the data type, unit, and data value are within the specified range. The completeness of fields in the data element information and whether the data conforms to the expected type and range are checked by comparing it with the format rules. The results of the format consistency verification are recorded as format identification information.
[0028] S1.2. Based on the format identification information, analyze the source, physical meaning and data fields of each data element, obtain the data type label, and perform format conversion on the data type label to generate an intermediate dataset.
[0029] It should be explained that the data collection method and acquisition channel are confirmed by identifying the source of each data element. Then, the physical meaning of each data element is further explored to understand its actual representation and role in the GIM model, clarifying the meaning and purpose of the data fields. Based on the characteristics of the data fields, the corresponding data type label for each data element is obtained, ensuring that the data label matches its physical meaning. The data type labels are then formatted by converting different types of data (such as integers, floating values, and dates) into a unified standard format and standardizing the units of the fields to ensure data consistency and generate an intermediate dataset.
[0030] S1.3 Perform integrity checks on the intermediate dataset and generate integrity verification information.
[0031] It should be noted that each data item in the intermediate dataset should be verified for missing data to ensure that all necessary data fields are included. The cause and location of data missing data should be traced and located using log files and error reports. Based on the information in the error reports, appropriate remedial measures should be taken to fill in or correct the missing data, ensuring data integrity. For example, missing numerical data can be filled using mean imputation and interpolation methods. After the repair is completed, integrity verification information should be generated.
[0032] S1.4 Check and repair missing entries in the integrity verification information to generate the original dataset.
[0033] It should be noted that each data entry should be checked for missing fields or null values. Missing entries and fields should be identified by comparing each record in the dataset with the expected data structure. The source of the missing data and the repair method should be determined by comparing integrity verification information with previous records and data sources. According to the repair strategy, the missing data should be filled in to ensure the integrity of the dataset. After the repair is complete, the original dataset should be generated.
[0034] It should also be noted that imputation strategies refer to the specific methods used to repair missing data. Common imputation strategies include: using the mean and most frequent values; for time series data, interpolation methods (such as linear interpolation and spline interpolation) can be used to infer missing values; for categorical data, the most common categories can be used for imputation.
[0035] S2. Remove invalid data from the original dataset and convert it into a unified standard format to generate a standardized dataset.
[0036] S2.1 Based on the original dataset, check whether the core fields of each data element are complete, whether the data values are within the valid range, and whether the format conforms to the standard, and remove invalid data to generate a cleaned dataset.
[0037] It should be noted that the core fields of each data meta-information should be checked to confirm that each field is complete and that all necessary fields are included without omission. The data values in each data meta-information should be checked to ensure they are within the valid range and meet actual requirements. The format of each data meta-information should be verified to ensure it conforms to standards. For data meta-information that does not meet the above conditions, invalid data should be removed, and a cleaned dataset should be generated.
[0038] S2.2. Convert the format of the cleaned dataset and standardize the intermediate dataset.
[0039] It should be noted that each data element in the cleaned dataset is reviewed individually to determine if its format conforms to the standard. According to the conversion rules, the data elements are converted from their existing format to a unified standard format to ensure data consistency and comparability. After format conversion, the converted data undergoes standardization processing to unify the data to the specified standard range and units, eliminating deviations and differences between data, ensuring data consistency and standardization, and generating a standardized intermediate dataset.
[0040] It should also be noted that conversion rules refer to the specifications and operations used in data processing to convert data from one format or type to another. Common conversion rules include data type conversion, unit unification, and field renaming.
[0041] S2.3. Perform integrity verification on the standardized intermediate dataset, repair missing data, and generate a standardized dataset.
[0042] It should be noted that each data element in the standardized intermediate dataset should be examined to verify the existence of any missing fields. Based on integrity criteria, the location and cause of missing data should be analyzed, and repair methods should be determined. For missing data, interpolation and imputation strategies should be used to fill in the missing parts, ensuring the integrity of the dataset. After repair, verification should be performed again to ensure that all missing data has been successfully filled, ultimately generating a complete and missing standardized dataset.
[0043] S3. Using dynamic mapping rules, the power transmission and transformation engineering facilities and structural components in the standardized dataset are mapped to the corresponding entities in the target format to generate a mapping dataset.
[0044] S3.1. Using dynamic mapping rules, each element in the standardized dataset is matched with the corresponding entity in the target format, and a mapping relationship is established to generate preliminary mapping rules.
[0045] It should be noted that each element in the standardized dataset is examined individually to ensure its structure and attributes are complete and matched against the corresponding entities in the target format (such as device model, sensor type, and acquisition time). Based on the type and attributes of the data elements in the standardized dataset, the most relevant entities in the target format are found; for example, the temperature field is matched against temperature sensor data in the target format. Matching is performed based on similarity and logical relationships. A mapping relationship between each element and the entity in the target format is established, clarifying how data is transformed from the standardized dataset to the target format, and generating preliminary mapping rules.
[0046] It should also be noted that similarity and logical relationship refer to ensuring that the matching between source data and target data conforms to predetermined rules and structures during the data mapping process. Similarity usually refers to the degree of similarity between data elements in terms of type, unit, and range. For example, temperature and humidity data may have similar units and can be matched based on similarity.
[0047] Logical relationships refer to the inherent connections and dependencies between data elements, such as the logical relationship between device model and sensor type. The device model determines the types of sensors it supports, and this relationship needs to be ensured to be correct in data mapping.
[0048] S3.2. Based on the preliminary mapping rules, analyze the specific conversion relationship between each data field in the standardized dataset and the entity attributes in the target format, and generate data mapping relationships.
[0049] It should be explained that, through the various data fields in the standardized dataset, the specific conversion relationship between each field and the entity attribute in the target format is determined. By comparing the definitions and type characteristics of the fields in the standardized dataset and the entity attributes in the target format, the mapping method ensures matching in terms of data type, unit, and precision. For example, the temperature value (in degrees Celsius) in the source data is correctly mapped to the temperature field in the target format, and its unit is converted from Fahrenheit to Celsius, while ensuring that the numerical precision is not lost. A detailed record is made of how each field is converted to the corresponding entity attribute in the target format, ensuring that the conversion relationship is clear and conforms to the requirements of the target format, thus generating a data mapping relationship.
[0050] S3.3 Based on the data mapping relationship, dynamically adjust the matching rules for field names and attribute types to generate adjusted mapping rules.
[0051] It should be explained that the matching between each field name in the standardized dataset and the entity attributes in the target format is verified. The naming rules of the data field names and the entity attributes in the target format are compared to identify any inconsistencies. Based on the actual needs of the data types and attributes, the matching rules between field names and attribute types are dynamically adjusted. For example, the voltage value field in the source data is adjusted to voltage and converted to a floating value type according to the target format requirements, while ensuring that the unit is uniformly converted from volts to kilovolts. This ensures that each field name is consistent with and logically reasonable in relation to the entity attribute type in the target format, generating adjusted mapping rules.
[0052] S3.4. Based on the adjusted mapping rules, map each element in the standardized dataset to the corresponding entity in the target format to generate a preliminary mapping dataset.
[0053] It should be noted that the field names of each data element should be checked to ensure they match the names of the corresponding entity attributes in the target format. The attribute types of each data element should be analyzed to ensure they match the types of entity attributes in the target format, such as data types (e.g., integers, floating values) and units (e.g., meters, kilowatts). This step-by-step comparative analysis ensures that each data element in the standardized dataset correctly matches the entity in the target format. Based on the adjusted mapping rules, each data element in the standardized dataset is converted into the form of its corresponding entity in the target format, ensuring that each data item is correctly mapped to the target entity. Through this mapping, all data elements are transformed into the structure and type required by the target format, generating a preliminary mapped dataset.
[0054] S3.5 Based on the initial mapping dataset, adjust the weights of each data field during the mapping process and output the mapping weight adjustment data.
[0055] It should be noted that the importance of each data field in the mapping process is evaluated to determine the contribution of each field to the mapping result of the target entity. Based on the evaluation results, the weights of each data field are adjusted to ensure that fields with a greater impact on the final mapping result receive higher weights, while fields with a smaller impact receive correspondingly lower weights. In this way, the weights of the data fields in the mapping process are adjusted to optimize the mapping accuracy, and the mapping weight adjustment data is output.
[0056] It should be noted that the expression for determining the contribution of each field to the target entity mapping result is as follows: ; in: For each field Contribution to the target entity mapping result; For fields In the target entity The conditional probability under given conditions is obtained through statistical estimation of data. For fields The total probability; It is the first in the dataset One data field; Indexes to data fields; For the target entity or target variable.
[0057] S3.6 Adjust the data according to the mapping weights, map each entity in the standardized dataset according to the target format requirements, and perform type conversion and unit unification to generate a mapped dataset.
[0058] It should be noted that the field names and attribute types of the entities should be checked to ensure they match the corresponding entity names and attribute types in the target format. The field values of each entity should be verified to conform to the range and units required by the target format. By comparing the entities in the standardized dataset with those in the target format, the mapping relationship of each entity should be ensured to be accurate and to meet the specifications and structural requirements of the target format. According to the requirements of the target format, necessary type conversions should be performed on each entity, such as converting numeric data to the appropriate unit type, to ensure consistency with the target format. Units in the data should be standardized to conform to the target format standard; for example, all temperature data should be standardized to degrees Celsius. After completing the mapping and unit conversion, the mapped dataset should be generated.
[0059] S4. Using machine learning algorithms, adjust the mapping relationship and data structure of the mapping dataset to obtain the optimized mapping dataset. Then, perform format compatibility checks and error repairs on the optimized mapping dataset to generate a verification dataset.
[0060] It should be noted that existing methods rely on manual or rule-driven approaches for data mapping and transformation. After manually setting the mapping rules, errors in the data are manually checked and corrected. This method is inefficient, susceptible to human factors, and struggles to handle complex data structures and large-scale datasets.
[0061] This invention uses machine learning algorithms to automatically adjust the mapping relationships and data structure of a mapping dataset. By training a model, it automatically optimizes the mapping process, reducing manual intervention. The optimized mapping dataset undergoes format compatibility checks and error correction to ensure data structure compatibility and consistency, thereby improving the efficiency and accuracy of data processing.
[0062] S4.1. Using the decision tree algorithm, type features are extracted, normalized, and one-hot encoded in the mapping dataset to convert the numerical and categorical fields in the mapping dataset into a unified format, generating a transformed dataset.
[0063] It should be noted that key attributes or variables are extracted from each data entry in the original dataset or a preprocessed dataset. For example, in a sensor dataset containing information such as temperature, humidity, and time, there are numerical and categorical fields. The correlation between each type of feature and the target variable is calculated using mutual information methods. For each type of feature, its contribution to the GIM model is evaluated based on its influence on the target variable. For numerical fields, normalization is performed to scale the data range to a uniform standard interval (e.g., 0 to 1) to ensure that data at different scales can be compared and processed under the same standard. Normalization is also performed on the numerical fields in the mapping dataset to transform data with different numerical ranges to a uniform standard range, ensuring data comparability and consistency. One-hot encoding is applied to the categorical fields in the mapping dataset to convert each categorical variable into an independent binary feature for subsequent processing. The data after feature extraction, normalization, and one-hot encoding generates a transformed dataset.
[0064] It should also be noted that information gain is a commonly used metric in feature selection, primarily used to measure the importance of a feature in data classification. A higher information gain indicates that the feature can better divide the dataset into different categories, improving classification accuracy.
[0065] S4.2 Adjust the feature weights and data structure in the transformed dataset, obtain the adjusted transformed dataset, and predict the mapping relationship between data fields in the adjusted transformed dataset to generate the optimized mapping dataset.
[0066] It should be noted that, based on the importance of each feature to the target variable, the weights of each parameter feature are reassessed to ensure that more parameter features have a greater weight in the dataset. Based on the adjusted weights of each parameter feature, the data structure is optimized, potentially involving feature selection and merging, to improve the quality and structural consistency of the dataset. After completing the data structure adjustment, the mapping relationships between data fields are predicted, the correlation between each field and other fields is analyzed, and the mapping relationships are corrected based on the prediction results to ensure more accurate matching between data fields, generating the optimized mapped dataset.
[0067] S4.3. Transform the table structure in the optimized mapping dataset into a graph structure, obtain the updated data structure, and perform a format compatibility check on the updated data structure to generate format verification data.
[0068] It should be noted that, based on the relationships between fields and entities in the mapped dataset, a relationship graph of nodes and edges is constructed. Each node represents an entity field in the dataset, and each edge represents the association between them. After forming the graph structure, the updated data structure is checked to ensure that the nodes and edges in the graph accurately reflect the true relationships between the data. A format compatibility check is performed on the updated data structure to ensure that it meets the requirements of the target format, and the consistency of node types, edge types, and association relationships is verified, generating format verification data.
[0069] S4.4 Perform field data repair and unit conversion on the format validation data to generate a validation dataset.
[0070] It should be explained that the data in each field should be checked to ensure it conforms to the expected format, and any mismatched fields should be identified. A detailed inspection of the data should be performed to identify any discrepancies in field names and data types. Unit conversion should be performed, addressing different units in the data to ensure all numeric fields conform to the target format requirements. After unit conversion, the data type and unit of each field should be verified again to ensure consistency and standardization. After completing field data repair and unit conversion, a verification dataset should be generated.
[0071] S5. Based on the validation dataset, compare the differences between the source data and the target data in terms of geometry, attribute data, and topology. Use a format compatibility verification tool to gradually detect and fix inconsistencies, and generate a conversion report.
[0072] It should be noted that this method involves manually comparing the differences in geometry, attributes, and topology between the source and target data one by one. After identifying the differences, manual intervention is required for correction, which is a tedious and error-prone process that lacks automation and flexibility, and is particularly inefficient when dealing with complex data.
[0073] This invention automatically detects differences between source and target data in terms of geometry, attribute data, and topology using a format compatibility verification tool, and gradually resolves inconsistencies through an automated repair process. This process significantly improves repair efficiency and accuracy, and ensures the completeness and accuracy of the conversion report content, avoiding errors and workload caused by manual intervention.
[0074] S5.1 Extract the geometric information, attribute data, and topological structure of the source and target data from the validation dataset, and use a geometric matching algorithm to compare the differences one by one to generate a difference comparison report.
[0075] It should be noted that the geometric information, attribute data, and topological structure of the source and target data are separated to ensure the accuracy and completeness of each data type. Geometric matching algorithms are used to calculate the similarity between the source and target data to determine their correspondence. Common geometric matching algorithms include point-to-point matching, boundary-to-boundary matching, and shape feature-based matching (such as angles, distances, and areas). By extracting geometric information (such as point coordinates, shape, and size) from the source and target data, algorithms (such as Euclidean distance calculation) are used to compare corresponding geometric elements one by one, detecting differences in position, shape, and size. For positional differences, adjustments are made through coordinate transformation; for shape differences, geometric transformations are used for correction; and for size differences, scaling is performed, generating a difference report.
[0076] It should be noted that the expression for calculating the similarity between the source and target geometric objects is: ; in: The similarity between the source and target geometric objects; As the source With the target point The Euclidean distance between them is used to measure the difference in their positions; Represents the coordinates of the source point; Represents the coordinates of the target point; Let be a constant, representing the maximum possible value of the Euclidean distance; Source geometry With target geometry Similarity measure between them; Features representing the geometry of the source; Features representing the geometry of the target; The maximum possible value of geometric shape similarity is determined by the maximum value of the defined geometric shape similarity metric, representing the similarity when there is a perfect match; The scale factor between the source and target geometric objects is used to measure the size difference and is determined by calculating the ratio of their sizes. This is the standard scaling factor, used to normalize scaling differences, and is determined by the data range.
[0077] S5.2 Based on the difference comparison report, analyze the format issues between the source data and the target data, and perform field type, unit and precision checks to generate format compatibility verification data.
[0078] It should be explained that the difference comparison report should be reviewed item by item to identify differences between the source and target data in field names, data types, and units. For each difference, the specific content of the source and target data should be examined to determine the cause of the difference, such as whether field names are consistent, data types match, or units conform to standards. By comparing the specific details of the source and target data, the nature of each difference should be analyzed to clarify the source and specific manifestation of formatting issues. The type of each field should be checked one by one to confirm whether the field types are consistent between the source and target data, analyze whether there are any type mismatches, and check the units to ensure that the units of numeric fields in the source and target data are consistent to avoid inconsistencies. The precision should be checked to ensure that the source and target data meet the precision requirements of the target format. Through these checks, format compatibility verification data should be generated.
[0079] S5.3 Verify the data based on format compatibility, detect geometric anomalies, adjust the data position and size, and generate a geometric repair dataset.
[0080] It should be noted that each part of the geometry is inspected to identify anomalies in size and location. For the detected anomalies, the sources of error in the geometry are adjusted to ensure that the data's position and size conform to the target format requirements. During the adjustment process, the coordinate position of the geometry is corrected to ensure data consistency, and the dimensions are appropriately scaled to generate a geometry-repaired dataset.
[0081] S5.4 Based on the geometric repair dataset, handle the mismatch of attribute data, unify the data format and units, and generate the attribute repair dataset.
[0082] It should be noted that each attribute data point is individually examined to identify any mismatches between the attribute data and the target format, determining which data fields exhibit inconsistencies. For each identified mismatch, the attribute data is corrected to ensure that the value and type of each field conforms to the target format requirements. The format of the attribute data is standardized to ensure data type consistency; for example, all numeric data is converted to a uniform format, and units are converted to ensure all units conform to the target standard. After these corrections are completed, an attribute repair dataset is generated.
[0083] S5.5 Repair the dataset based on attributes, resolve topology defects, adjust device connection relationships, repair missing topology connections, and generate a topology repair dataset.
[0084] It should be noted that the process involves checking the connections between each node to ensure that the connection information of each node is consistent with the topology, and identifying any inconsistencies in the connections, such as incorrect connection paths or missing connections. The topology is checked for compliance with rules to ensure that there are no missing connections and that the connections between each node are complete and accurate. Defects are identified and located by comparing the existing topology with the target structure. Based on the connection relationships, it is determined which connections need adjustment to ensure that the connections between devices meet the requirements of the target format, and missing topology connections are repaired to ensure the completeness of the connection relationships between each device, correcting any errors or omissions in the topology. After the adjustments are completed, a topology repair dataset is generated.
[0085] S5.6. Based on the topology repair dataset, record the repair process of geometry, attribute data, and topology, and generate a conversion report.
[0086] It should be noted that every step of the geometry, attribute data, and topology repair process should be meticulously documented, ensuring complete record-keeping for each repair operation. The specific measures taken during the repair process should be described item by item, including adjustments to the geometry's position, standardization of attribute data format and unit conversion, and repair of missing connections in the topology. A detailed process report should be compiled for all repair steps, ensuring each repair operation is clear and traceable, and a transformation report should be generated.
[0087] In summary, this invention achieves the following: by collecting and identifying GIM model data formats, removing invalid data and converting it to a standard format, optimizing data mapping using dynamic mapping rules and machine learning, and performing format compatibility checks and error correction; industrial data processing plays a crucial role in this process, ensuring the accuracy and consistency of data conversion; by comparing and correcting differences between source and target data in geometry, attributes, and topology, it greatly improves the efficiency and accuracy of data processing, providing an efficient and flexible automated data conversion solution.
[0088] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for automatic multi-format conversion of GIM models for power transmission and transformation engineering, characterized in that: include, Collect GIM model data, identify the format of the source data in the GIM model data and mark the data type, and generate the original dataset; Remove invalid data from the original dataset and convert it into a unified standard format to generate a standardized dataset; Dynamic mapping rules are used to map power transmission and transformation engineering facilities and structural components in the standardized dataset to corresponding entities in the target format, generating a mapping dataset; By using machine learning algorithms, the mapping relationships and data structures of the mapping dataset are adjusted to obtain an optimized mapping dataset. The optimized mapping dataset is then subjected to format compatibility checks and error correction to generate a validation dataset. Based on the validation dataset, the differences between the source data and the target data in terms of geometry, attribute data, and topology are compared. Inconsistencies are detected and corrected step by step using a format compatibility verification tool, and a conversion report is generated.
2. The automatic multi-format conversion method for GIM models in power transmission and transformation engineering as described in claim 1, characterized in that: The specific steps for collecting GIM model data, identifying the format of the source data in the GIM model data, and marking the data type are as follows. Parse the data metadata of GIM model data, identify and verify data format consistency, and generate format identification information; Based on the format recognition information, analyze the source, physical meaning and data fields of each data element, obtain the data type label, and perform format conversion on the data type label to generate an intermediate dataset; Perform integrity checks on the intermediate dataset and generate integrity verification information.
3. The automatic multi-format conversion method for GIM models for power transmission and transformation projects as described in claim 2, characterized in that: The process of generating the original dataset refers to checking and repairing missing entries in the integrity verification information.
4. The automatic multi-format conversion method for GIM models for power transmission and transformation projects as described in claim 3, characterized in that: The process of removing invalid data from the original dataset and converting it into a unified standard format to generate a standardized dataset involves the following steps. Based on the original dataset, check whether the core fields of each data element are complete, whether the data values are within the valid range, and whether the format meets the standards, and remove invalid data to generate a clean dataset; Perform format conversion on the cleaned dataset and standardize the intermediate dataset; Perform integrity verification on the standardized intermediate dataset, repair missing data, and generate a standardized dataset.
5. The automatic multi-format conversion method for GIM models for power transmission and transformation projects as described in claim 4, characterized in that: The method employs dynamic mapping rules to map power transmission and transformation engineering facilities and structural components in the standardized dataset to corresponding entities in the target format. The specific steps are as follows. By employing dynamic mapping rules, each element in the standardized dataset is matched with the corresponding entity in the target format, and a mapping relationship is established to generate preliminary mapping rules. Based on the preliminary mapping rules, analyze the specific conversion relationship between each data field in the standardized dataset and the entity attributes in the target format, and generate data mapping relationships; Based on the data mapping relationship, the matching rules for field names and attribute types are dynamically adjusted to generate the adjusted mapping rules; Based on the adjusted mapping rules, each element in the standardized dataset is mapped to the corresponding entity in the target format to generate a preliminary mapping dataset; Based on the initial mapping dataset, the weights of each data field during the mapping process are adjusted, and the mapping weight adjustment data is output.
6. The automatic multi-format conversion method for GIM models for power transmission and transformation projects as described in claim 5, characterized in that: The generated mapping dataset refers to adjusting the data according to the mapping weights, mapping each entity in the standardized dataset according to the target format requirements, and performing type conversion and unit unification.
7. The automatic multi-format conversion method for GIM models for power transmission and transformation projects as described in claim 6, characterized in that: The process of adjusting the mapping relationships and data structure of the mapping dataset using machine learning algorithms to obtain an optimized mapping dataset involves the following specific steps. The decision tree algorithm is used to extract features, normalize and encode one-hot data in the mapping dataset, and convert the numerical and categorical fields in the mapping dataset into a unified format to generate a converted dataset. Adjust the feature weights and data structure in the transformed dataset to obtain the adjusted transformed dataset, and predict the mapping relationship between data fields in the adjusted transformed dataset to generate the optimized mapped dataset.
8. The automatic multi-format conversion method for GIM models for power transmission and transformation projects as described in claim 7, characterized in that: The optimized mapping dataset undergoes format compatibility checks and error correction to generate a validation dataset. The specific steps are as follows: The table structure in the optimized mapping dataset is transformed into a graph structure, the updated data structure is obtained, and the format compatibility of the updated data structure is checked to generate format verification data. Perform field data repair and unit conversion on the format validation data to generate a validation dataset.
9. The automatic multi-format conversion method for GIM models in power transmission and transformation engineering as described in claim 8, characterized in that: The process involves comparing the source and target data based on the validation dataset, identifying and correcting inconsistencies through a format compatibility checker. The specific steps are as follows: Geometric information, attribute data, and topological structure of source and target data are extracted from the validation dataset, and a geometric matching algorithm is used to compare the differences one by one to generate a difference comparison report. Based on the difference comparison report, analyze the format issues between the source data and the target data, and perform field type, unit and precision checks to generate format compatibility verification data; Based on the format compatibility verification data, geometric anomalies are detected, and the data position and size are adjusted to generate a geometric repair dataset; Based on the geometric repair dataset, the mismatch of attribute data is handled, and the data format and units are unified to generate the attribute repair dataset; Based on the attribute repair dataset, the defects in the topology structure are resolved, the device connection relationships are adjusted, the missing topology connections are repaired, and a topology repair dataset is generated.
10. The automatic multi-format conversion method for GIM models for power transmission and transformation projects as described in claim 9, characterized in that: The generation of the transformation report refers to the process of recording the repair of geometry, attribute data, and topology based on the topology repair dataset.