An AI-based method for extracting technical indicators from transmission line engineering data

By screening significant technical indicators in transmission line engineering data through AI-based methods, the problem that traditional methods are difficult to obtain accurately is solved, and more efficient engineering data utilization and decision-making support are achieved.

CN120087913BActive Publication Date: 2025-08-12STATE GRID JIANGXI ELECTRIC POWER CO LTD ECONOMIC & TECH RES INST +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510151893.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-08-12
Estimated Expiration
2045-02-11

AI Technical Summary

Technical Problem

Traditional technical indicator extraction methods are difficult to quickly and accurately obtain technical indicators in transmission line engineering data, and cannot meet the technical requirements of different construction conditions, resulting in low utilization rate of engineering data.

Method used

Using an AI-based method, by obtaining the histogram of technical indicators, constructing a rightless undirected graph, analyzing the cluster clusters and construction contribution weights, filtering significant general and scenario-based technical indicators, classifying and clustering, and improving the screening accuracy.

Benefits of technology

It greatly reduces the number of general technical indicators, reduces the dimensions of subsequent cluster analysis and big data mining, improves the utilization rate of transmission line engineering data, and provides more accurate decision-making support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120087913B_ABST
    Figure CN120087913B_ABST
Patent Text Reader

Abstract

The present application relates to the field of data mining and processing technology, and specifically to an AI-based method for extracting technical indicators from transmission line engineering data, including: obtaining technical indicators contained in each transmission line engineering data, determining the aggregation characteristic value of each technical indicator based on the degree of data distribution aggregation around the peak in the histogram of each technical indicator, constructing a weighted undirected graph to determine the significant screening value of each technical indicator, dividing significant general technical indicators and non-general technical indicators, determining the neighboring expansion scale of each non-general technical indicator in each cluster, analyzing the numerical differences of non-general technical indicators in the transmission line engineering data within the cluster, determining the construction contribution weight, obtaining significant scenario-based technical indicators, and completing the extraction of technical indicators from the transmission line engineering data. The present application can improve the utilization rate of transmission line engineering data and ensure the effectiveness of technical indicators.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data mining and processing technology, and specifically to a method for extracting technical indicators of transmission line engineering data based on AI. Background Art

[0002] Transmission line projects primarily involve three phases: project planning and design, construction, and maintenance. Each phase generates a significant amount of transmission line engineering data, which contains numerous core technical indicators reflecting the safety performance of the transmission line. Extracting technical indicators from this data ensures that construction strictly adheres to design specifications and drawings, allowing for the timely identification and correction of quality issues. This also helps to appropriately reduce project costs and effectively control them.

[0003] Traditional technical indicator extraction methods typically use natural language processing and large language models to directly extract technical indicators from transmission line engineering data to provide data support for transmission line construction. However, engineering data contains a vast number of technical indicators, and transmission lines under different construction conditions have different technical requirements for these indicators. Traditional technical indicator extraction methods fail to screen these indicators, making it increasingly difficult to meet the industry's demand for rapid and accurate technical indicator acquisition. Therefore, there is an urgent need to mine, screen, and efficiently utilize these technical indicators. Summary of the Invention

[0004] In order to solve the above technical problems, this application provides an AI-based method for extracting technical indicators of transmission line engineering data to solve existing problems.

[0005] The present application's AI-based method for extracting technical indicators from transmission line engineering data adopts the following technical solutions:

[0006] One embodiment of the present application provides an AI-based method for extracting technical indicators from transmission line engineering data, comprising the following steps:

[0007] Obtain technical indicators contained in the engineering data of each transmission line;

[0008] Analyze the distribution of each technical indicator and construct a histogram of each technical indicator. According to the degree of data distribution aggregation around the peak in the histogram of each technical indicator, determine the aggregation characteristic value of each technical indicator.

[0009] Analyze the correlation between any two technical indicators on the histogram to construct a weighted undirected graph, analyze the node score of each technical indicator in the weighted undirected graph, combine the degree of each technical indicator corresponding to the node in the weighted undirected graph, determine the significant screening value of each technical indicator, and classify the significant general technical indicators and non-general technical indicators;

[0010] The transmission line engineering data are classified based on the significant universal technical indicators in each transmission line engineering data. Based on the inter-cluster distance between each cluster and other clusters and the analysis of the mutation of the inter-cluster distance, combined with the significant screening values of each non-universal technical indicator, the neighboring expansion scale of each non-universal technical indicator of each cluster is determined.

[0011] Analyze the numerical differences of non-general technical indicators in the transmission line engineering data within the clusters, determine the construction contribution weights of each non-general technical indicator in each cluster, and obtain significant scenario-based technical indicators. The significant general technical indicators and significant scenario-based technical indicators are used as technical indicators extracted from the transmission line engineering data.

[0012] Preferably, the construction of the histogram of each technical indicator further includes:

[0013] For each technical indicator, the values of the technical indicators of all transmission lines are normalized, and the normalized results are used as the horizontal axis in ascending order, and the frequency of occurrence of the normalized technical indicator values is used as the vertical axis to obtain the histogram of each technical indicator.

[0014] Preferably, the calculation method of the aggregate characteristic value of each technical indicator is:

[0015] Where, is the aggregate eigenvalue of the technical indicator q, is the number of peak points in the histogram of the technical indicator q, is the sum of all elements in the growth set corresponding to the v-th peak point, is the clustering interval distance of the vth peak point, which is obtained by calculating the distance between the vth peak point and its nearest peak point. D is the mean of the clustering interval distances of all peak points, and exp() is an exponential function with a natural constant as the base.

[0016] Preferably, the construction of the weighted undirected graph further includes:

[0017] Calculate the absolute value of the Pearson correlation coefficient between any two technical indicator histogram sequences as the correlation value of the any two technical indicators, perform threshold segmentation on the correlation values between all any two technical indicators to obtain a first segmentation threshold, and record the technical indicator combination with a correlation value greater than the first segmentation threshold as having a strong correlation;

[0018] All technical indicators are regarded as nodes, and two technical indicators with strong correlation are connected. The edge weight of the connection is the correlation value of the technical indicator combination. A self-loop is constructed for each technical indicator node, and the edge weight of the self-loop is the aggregation eigenvalue to obtain a weighted undirected graph of all technical indicators.

[0019] Preferably, the significant screening value of each technical indicator is a normalized result of the product of the node score of each technical indicator in the weighted undirected graph and the degree of the node corresponding to each technical indicator in the weighted undirected graph.

[0020] Preferably, the classification of significant general technical indicators and non-general technical indicators includes:

[0021] The significant screening values of all technical indicators are threshold segmented to output a second segmentation threshold, and the technical indicators with significant screening values greater than the second segmentation threshold are regarded as significant general technical indicators, and all other technical indicators are regarded as non-general technical indicators.

[0022] Preferably, the classification of the transmission line engineering data further includes:

[0023] The transmission line engineering data are taken as samples, and the normalized results of the numerical values of all significant universal technical indicators of each sample are used to form the universal indicator vector of each sample. A clustering algorithm is used to cluster all transmission line engineering data according to the universal indicator vectors of all samples to obtain clusters.

[0024] Preferably, the calculation method of the proximity expansion scale of each non-universal technical indicator of each cluster is:

[0025] Where, is the neighboring expansion scale of the xth non-universal technical indicator of the ith cluster, is the element value of the first mutation point in the inter-cluster distance sequence of the i-th cluster, is the significant screening value of the xth non-general technical indicator, where the inter-class distances between the i-th cluster and all other clusters are arranged in ascending order of values to form the inter-class distance sequence of the i-th cluster.

[0026] Preferably, determining the construction contribution weight of each non-general technical indicator of each cluster to obtain a significant scenario-based technical indicator further includes:

[0027] Extracting adjacent clusters of each cluster through the relationship between the inter-cluster distance and the neighbor expansion scale;

[0028] Taking the cluster center of any cluster as the target sample, calculate the distance between all samples in the cluster and all its neighboring clusters and the target sample, and sort them in ascending order of distance to extract L samples as the positive sample set; accordingly, calculate the distance between all samples in all clusters except the cluster and all its neighboring clusters and the target sample, and similarly sort them in ascending order of distance to extract L samples as the negative sample set;

[0029] The calculation method of the construction contribution weight of each non-general technical indicator of each cluster is as follows: Where, The construction contribution weight of the xth non-general technical indicator, 、 They are respectively the positive sample set and the negative sample set. The value of the xth non-general technical indicator for each sample, is the value of the target sample regarding the xth non-general technical indicator;

[0030] All calculated construction contribution weights are normalized, and non-general technical indicators whose normalized construction contribution weights are greater than or equal to the contribution threshold are regarded as significant scenario-based technical indicators.

[0031] Preferably, the extraction of adjacent clusters of each cluster further includes: for each cluster, taking clusters whose inter-cluster distance is less than or equal to the adjacent expansion scale as adjacent clusters of each cluster.

[0032] This application has at least the following beneficial effects:

[0033] This application obtains significant screening values for each technical indicator based on the standardized characteristics of transmission line projects. Its beneficial effect is to further screen significant general technical indicators in transmission line projects, significantly reduce the number of general technical indicators obtained through transmission line screening, and reduce the dimensions of subsequent cluster analysis and big data mining.

[0034] This application considers the problem of imbalanced categories of transmission line engineering data in various construction scenarios, adaptively obtains neighboring expansion scales, and adopts neighboring cluster samples with similar specifications and construction scenarios to solve the problem of scarcity of samples of rare construction scenarios; it conducts big data mining on the massive technical indicators of any transmission line engineering data, and accurately screens significant general technical indicators and significant scenario-based technical indicators for the construction scenarios of the transmission line engineering data as decision-making indicators for the transmission line engineering data, thereby improving the utilization rate of the transmission line engineering data, eliminating a large number of redundant technical indicators, and providing strong support for engineering construction, operation and maintenance, and other links. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present application or the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0036] Figure 1 A flowchart of the steps of an AI-based method for extracting technical indicators of transmission line engineering data provided in this application;

[0037] Figure 2 A schematic diagram of the construction process of the weighted undirected graph provided in this application. DETAILED DESCRIPTION

[0038] In order to further illustrate the technical means and effects adopted by this application to achieve the predetermined invention objectives, the following, in conjunction with the accompanying drawings and preferred embodiments, describes in detail the specific implementation, structure, features, and effects of an AI-based method for extracting technical indicators of transmission line engineering data proposed in this application. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures, or characteristics of one or more embodiments may be combined in any suitable form.

[0039] Unless otherwise defined, terms such as "comprises," "comprising," or any other variants thereof are intended to encompass non-exclusive inclusion, such that a circuit structure, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such article or device. In the absence of further restrictions, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the article or device comprising the element. In addition, the term "and\or" as used herein includes any and all combinations of one or more related listed items. All technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application pertains.

[0040] The following describes in detail a specific method for extracting technical indicators of transmission line engineering data based on AI provided by this application with reference to the accompanying drawings.

[0041] An embodiment of the present application provides an AI-based method for extracting technical indicators from transmission line engineering data. For details, please refer to Figure 1 , including the following steps:

[0042] Step 1: Obtain the technical indicators contained in the engineering data of each transmission line.

[0043] Based on the characteristics of the transmission line project data, paper engineering documents were uniformly converted into Word format using OCR technology. Abbyy FineReader software was used for document scanning and format conversion. Text annotation tools were used to annotate the Word engineering documents. These were categorized into three types of labels: indicator name (Indicator), indicator value (Value), and indicator unit (Unit). Label Studio and YEDDA were used as text annotation tools.

[0044] In this embodiment, a bidirectional LSTM (Bidirectional Long ShortTerm Memory Network, BiLSTM) and a conditional random field (CRF) are used to establish an indicator extraction model to extract indicators from Word-formatted engineering documents. Specifically, the bidirectional LSTM is used to obtain contextual features of engineering document sentences and technical indicators. An attention mechanism is integrated into the indicator extraction model to emphasize important information in the contextual features. The outputs of the bidirectional LSTM layer and the attention layer are combined in a fully connected layer and used as input to the conditional random field layer to obtain indicator items for the transmission line engineering documents. It should be noted that in this embodiment, the indicator extraction model is trained using SGD stochastic gradient descent as the optimization algorithm, with the SGD momentum set to 0.9 and the loss function using binary cross entropy. The training of the indicator extraction model is well-known technology, and the specific process is not detailed here. In actual application scenarios, implementers may also use other methods to extract technical indicators from transmission line engineering documents, and this embodiment does not impose any specific restrictions on this.

[0045] In this embodiment, the technical indicators in the transmission line engineering data include each transmission line's line voltage, foundation compressive strength, foundation cross-sectional dimensions, tower structure inclination, horizontal span, conductor sag, distance to ground, transmission power, transmission capacity, maximum wind speed, maximum temperature, minimum temperature, average temperature, number of lightning days, number of lightning hours, maximum ice thickness, soil freezing depth, soil resistivity, and ground resistance. The foundation compressive strength refers to the compressive strength of the cast-in-place concrete or precast concrete components of the transmission line foundation; the horizontal span refers to the arithmetic average of the spans on both sides of the transmission line tower; the distance to ground refers to the minimum distance between any live part of the transmission line and the ground; and the transmission power refers to the maximum power the line can transmit.

[0046] Step 2: Analyze the distribution of each technical indicator and construct a histogram of each technical indicator. According to the degree of data distribution aggregation around the peak in the histogram of each technical indicator, determine the aggregation characteristic value of each technical indicator.

[0047] In the field of transmission line engineering, to improve design efficiency and quality, projects often adopt similar specifications based on established design standards and specifications. For example, for transmission lines with a 220kV voltage and a transmission capacity of 500MW, standards such as electrical safety distances and insulation coordination are referenced during design, ensuring similar specifications across transmission line projects. Technical indicators are primarily categorized as general and scenario-based. General technical indicators are well-suited to transmission line projects of varying specifications and construction conditions and serve as essential engineering indicators. Scenario-based indicators, on the other hand, provide effective decision support for transmission line construction under specific construction scenarios.

[0048] In this example, the numerical distribution of various technical indicators across all transmission line engineering data is analyzed. Min-Max normalization is used for normalization, and the normalized results are plotted in ascending order on the horizontal axis, with the vertical axis representing the frequency of occurrence of the normalized technical indicator values. A histogram of each technical indicator is constructed. The universal technical indicators in transmission line engineering data effectively reflect transmission line specifications. Due to the standardized nature of transmission line engineering, significant universal technical indicators are graded and exhibit distinct non-uniform clustering patterns on the histogram.

[0049] In this embodiment, a curve fitting is performed on the histogram of the technical indicator, local peak points are eliminated, and then all peak points of the fitted curve are obtained. Any peak point is used as the initial growth point, and the growth direction is that the initial center is on the left and right sides of the fitting curve. The growth criterion is that the value of the adjacent points is at least 1 / 2 of the value of the peak point. The region growing algorithm is used to obtain the growth set of each peak point. In this embodiment, the clustered characteristic value of any technical indicator is obtained using the following formula:

[0050] Where, is the aggregate eigenvalue of the technical indicator q, is the number of peak points in the histogram of the technical indicator q, is the sum of all elements in the growth set corresponding to the v-th peak point, is the clustering interval distance of the vth peak point, which is obtained by calculating the distance between the vth peak point and its nearest peak point. D is the mean of the clustering interval distances of all peak points, and exp() is an exponential function with a natural constant as the base.

[0051] In the above formula of this embodiment, the calculation The purpose of this is to prevent multiple peaks from occurring within a similar range of values for technical indicators. The larger the sum of all the values within the peak growth set, the more significant the clustering of the transmission line technical indicator, and the more qualified the technical indicator is to serve as a transmission line specification standard. Furthermore, the more discrete the distribution of peak points across all value ranges, the more likely it is that the technical indicator will have multiple levels of transmission line specification standards.

[0052] Step 3: Analyze the correlation between any two technical indicators on the histogram to construct a weighted undirected graph, analyze the node scores of each technical indicator in the weighted undirected graph, combine the degrees of the nodes corresponding to each technical indicator in the weighted undirected graph, determine the significant screening value of each technical indicator, and divide the significant general technical indicators into non-general technical indicators.

[0053] Transmission lines build a comprehensive engineering system through the combination of technical indicators, helping construction units strictly adhere to design specifications and drawings for standardized construction, and helping supervisors promptly identify and correct engineering quality issues. Therefore, in massive amounts of transmission line engineering data, two highly correlated technical indicators often appear in this combination, and the histogram distribution of these indicators exhibits extremely high correlation. Therefore, the similarity measurement results of any two technical indicators are analyzed. Specifically, in this embodiment, the absolute value of the Pearson correlation coefficient between any two technical indicator histogram sequences is calculated as the correlation value of the two technical indicators.

[0054] In this embodiment, all the correlation values between technical indicators are used as input to the maximum inter-class variance OSTU algorithm, which outputs a first segmentation threshold. Technical indicator combinations with correlation values greater than the first segmentation threshold are considered highly correlated. Universal technical indicators exhibit distinct non-uniformly spaced clustering characteristics on the histogram. For a single technical indicator, the higher the clustering characteristic value, the more likely it is a universal technical indicator. Considering the systematization and standardization of transmission line engineering, combinations of technical indicators with strong correlations are both highly likely to be universal technical indicators.

[0055] All technical indicators are taken as nodes, and technical indicators with strong correlation are connected. The edge weight of the connection is the correlation value of the technical indicator combination. A self-loop is constructed for each technical indicator node, and the edge weight of the self-loop is the aggregation feature value to obtain a weighted undirected graph of all technical indicators. In this embodiment, the flowchart of the construction of the weighted undirected graph is as follows: Figure 2 shown.

[0056] Furthermore, a weighted undirected graph is used as the input of the TextRank algorithm. The damping coefficient of the TextRank algorithm is set to 0.85, and the node score of each technical indicator is output. Technical indicators with higher node scores are considered to be generally important technical indicators in all construction scenarios of transmission line projects.

[0057] The correlation value between two technical indicators reflects the relationship between them. Universal indicators, as basic engineering indicators with strong adaptability, tend to have stronger correlations with more technical indicators than scenario-based technical indicators. In a weighted undirected graph of all technical indicators, the degree of each technical indicator node is obtained. The method for calculating the degree of a node in a weighted undirected graph is well known and will not be described in detail in this embodiment.

[0058] Furthermore, the product of the node score and degree of each technical indicator is normalized and used as the significant screening value of each technical indicator. The significant screening value is used to further screen the significant general technical indicators in the transmission line project, which can greatly reduce the number of general technical indicators obtained by transmission line screening and reduce the dimensions of subsequent cluster analysis and big data mining.

[0059] The significant screening values of all technical indicators are used as the input of the maximum inter-class variance OSTU algorithm, and the second segmentation threshold is output. The technical indicators with significant screening values greater than the second segmentation threshold are regarded as significant general technical indicators, and all technical indicators except the significant general technical indicators are recorded as non-general technical indicators.

[0060] Step 4: Classify the transmission line engineering data based on the significant universal technical indicators in each transmission line engineering data. Based on the inter-cluster distance between each cluster and other clusters, analyze the mutation of the inter-cluster distance, and combine the significant screening values of each non-universal technical indicator to determine the neighboring expansion scale of each non-universal technical indicator in each cluster.

[0061] In transmission line projects, due to the variability of construction scenarios, different specifications and standards are required in different geographical environments and climatic conditions, such as mountainous areas, plains, and coastal areas. These scenarios also place varying emphasis on scenario-based technical indicators. For example, in coastal areas, where the risk of typhoons is higher, transmission line projects focus more on wind speed indicators. In plain areas, where there are no obstructions such as high mountains, transmission line projects place greater emphasis on lightning and ground resistance indicators. To better facilitate the supervision of transmission line projects, it is necessary to extract scenario-based technical indicators from the vast number of technical indicators that meet technical requirements and provide more detailed construction information for each construction scenario.

[0062] The combination of all significant universal technical indicators can reflect the specifications, standards, and construction scenarios of a transmission line project. In this embodiment, each transmission line project data is used as a sample, and the normalized values of all significant universal technical indicators for each sample are combined to form a universal indicator vector for each sample. It should be noted that there are many methods for normalizing the values of each significant universal technical indicator. This embodiment uses the Min-Max normalization method. The specific normalization process is a well-known technique and is not described in detail in this embodiment. Furthermore, the universal indicator vectors of all samples are used as input to a clustering algorithm, which outputs a clustering result. It should be noted that all transmission line project data with highly similar specifications, standards, and construction scenarios are included in the same cluster. Specifically, the clustering algorithm can use the DPC (Density Peak Clustering) algorithm. The specific clustering process is a well-known technique and is not described in detail in this embodiment. In actual application scenarios, implementers may also choose other clustering algorithms, and this embodiment does not impose any specific restrictions on this.

[0063] There is an imbalance in the categories of transmission line engineering data in various construction scenarios. Some construction scenarios are common and have many construction cases, so there is a lot of corresponding transmission line engineering data. However, for those rare construction scenarios, transmission line engineering data is scarce. In the process of selecting technical indicators of transmission line engineering data, if only the construction scenario corresponding to each cluster is considered, it is difficult to accurately obtain technical indicators that meet industry requirements for construction scenarios with scarce data and a small number of samples. Samples and clusters with similar distribution distances in the feature space have similar specifications and construction scenarios. Therefore, in this embodiment, samples from adjacent clusters are used to make up for the scarcity of samples from rare construction scenarios. Therefore, in this embodiment, the inter-class distance between any two clusters will be obtained, where the calculation process of the inter-class distance is an existing technology and will not be repeated in this embodiment.

[0064] Specifically, the inter-class distance between the i-th cluster and the j-th cluster is recorded as The inter-class distances between the i-th cluster and all other clusters are arranged in ascending numerical order to form the inter-class distance sequence of the i-th cluster. This inter-class distance sequence is then used as the input for the BG sequence segmentation algorithm, and the first mutation point of the inter-class distance sequence is output. The larger the inter-class distance between the cluster corresponding to the mutation point and the i-th cluster, the less suitable it is to serve as a neighboring cluster of the i-th cluster. Therefore, this embodiment calculates the neighboring expansion scale of each cluster with respect to each non-universal technical indicator using the mutation point data in the inter-class distance sequence of the clusters and the significant screening values of each non-universal technical feature. The calculation formula in this embodiment is:

[0065] Where, is the neighboring expansion scale of the xth non-universal technical indicator of the ith cluster, is the element value of the first mutation point in the inter-cluster distance sequence of the i-th cluster, It is the significant screening value of the xth non-general technical indicator.

[0066] For technical indicators with low significant screening values, the higher the possibility that they belong to scenario-based technical indicators, the more obvious the problem of sample scarcity within the cluster. A larger neighboring expansion scale and more neighboring clusters are needed to compensate for the scarcity of samples from rare construction scenarios.

[0067] Step 5: Analyze the numerical differences of non-general technical indicators in the transmission line engineering data within the clusters, determine the construction contribution weights of each non-general technical indicator in each cluster, and obtain significant scenario-based technical indicators. The significant general technical indicators and significant scenario-based technical indicators are used as technical indicators extracted from the transmission line engineering data.

[0068] Furthermore, for the xth non-universal technical indicator of the i-th cluster, the inter-class distance is less than or equal to the adjacent expansion scale. The other clusters are regarded as the neighboring clusters of the i-th cluster.

[0069] Taking the cluster center of the i-th cluster as the target sample, calculate the distance between all samples and the target sample in the i-th cluster and all its neighboring clusters, and sequentially extract L samples in ascending order of distance as the positive sample set NearH. The samples in the positive sample set NearH have similar construction scenes as the target samples. The purpose of setting the number of samples L is to find representative samples of different construction scenes. In this embodiment, L is taken as 30. Correspondingly, calculate the distance between all samples and the target sample in all clusters except the i-th cluster and all its neighboring clusters. Similarly, arrange them in ascending order of distance and sequentially extract L samples as the negative sample set NearM. The samples in the negative sample set NearM have large differences in construction scenes with the target samples. It should be noted that in this embodiment, the above distance is measured using Euclidean distance, which can be selected by the implementer in actual application scenarios.

[0070] Furthermore, the value of the target sample with respect to the xth non-general technical indicator is recorded as , the positive sample set The value of the xth non-general technical indicator of the sample is recorded as , the negative sample set The value of the xth non-general technical indicator of the sample is recorded as , and then analyze the construction contribution weight of each non-general technical indicator of each cluster. The specific calculation formula is:

[0071] Where, The construction contribution weight of the xth non-general technical indicator.

[0072] It is understandable that the better the beneficial effect of the x-th non-general technical indicator on distinguishing transmission line construction scenarios, the more construction detail information can be captured through the technical indicator, and the higher the contribution to construction supervision, the more attention should be paid to the technical indicator with a high construction contribution weight for the i-th cluster.

[0073] At this point, the construction contribution weights of each cluster for each non-universal technical indicator are obtained. All construction contribution weights are normalized using the Max-Min method. Non-universal technical indicators with normalized construction contribution weights greater than or equal to the contribution threshold are retained and considered the significant scenario-based technical indicators for the cluster. In this example, the contribution threshold is set to 0.5.

[0074] Big data mining is conducted on the massive technical indicators of transmission line engineering data. According to the construction scenarios of transmission line engineering data, general technical indicators and scenario-based technical indicators are accurately screened as decision-making indicators for transmission line engineering data. This aims to improve the utilization rate of transmission line engineering data, eliminate a large number of redundant technical indicators, and provide strong support for engineering construction, operation and maintenance, and other links.

[0075] It is understood that references to "one embodiment" or "some embodiments" in the present specification mean that one or more embodiments of the present application include a particular feature, structure, or characteristic described in conjunction with that embodiment. Thus, if "in one embodiment," "in some embodiments," "in other embodiments," or "in other embodiments" appear in different places in this specification, they do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and their variations all mean "including but not limited to," unless otherwise specifically emphasized.

[0076] It should be noted that the above-mentioned sequence of the embodiments of the present application is for description only and does not represent the advantages and disadvantages of the embodiments. The above description is of a specific embodiment of this specification. In addition, the processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multi-tasking and parallel processing are also possible or may be advantageous. At the same time, the size of the sequence number of each step in the embodiment does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments in this specification.

[0077] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A method for extracting technical indicators of transmission line engineering data based on AI, characterized in that: The following steps are involved: Obtain technical indicators contained in the engineering data of each transmission line; For each technical indicator, the values of the technical indicators of all transmission lines are normalized. The normalized results are arranged in ascending order as the horizontal axis, and the frequency of occurrence of the normalized technical indicator values is used as the vertical axis to obtain the histogram of each technical indicator. The aggregation characteristic value of each technical indicator is determined based on the degree of data distribution aggregation around the peak value in the histogram of each technical indicator. Calculate the absolute value of the Pearson correlation coefficient between any two technical indicator histogram sequences as the correlation value of any two technical indicators. Based on the correlation value of the technical indicator combination, construct a weighted undirected graph of the technical indicators. Analyze the node scores of each technical indicator in the weighted undirected graph. Combined with the degree of each technical indicator's corresponding node in the weighted undirected graph, determine the significant screening value of each technical indicator and classify significant general technical indicators into non-general technical indicators. The transmission line engineering data are classified based on the significant universal technical indicators in each transmission line engineering data. Based on the inter-cluster distance between each cluster and other clusters and the analysis of the mutation of the inter-cluster distance, combined with the significant screening values of each non-universal technical indicator, the neighboring expansion scale of each non-universal technical indicator of each cluster is determined. Analyze the numerical differences of non-general technical indicators in the transmission line engineering data within the clusters, determine the construction contribution weights of each non-general technical indicator in each cluster, and obtain significant scenario-based technical indicators. The significant general technical indicators and significant scenario-based technical indicators are used as technical indicators extracted from the transmission line engineering data.

2. The method for extracting technical indicators of power transmission line engineering data based on AI according to claim 1, characterized in that: The calculation method of the aggregation characteristic value of each technical indicator is: Where, is the aggregate eigenvalue of the technical indicator q, is the number of peak points in the histogram of the technical indicator q, is the sum of all elements in the growth set corresponding to the v-th peak point, is the clustering interval distance of the vth peak point, which is obtained by calculating the distance between the vth peak point and its nearest peak point. D is the mean of the clustering interval distances of all peak points, and exp() is an exponential function with a natural constant as the base.

3. The AI-based method for extracting technical indicators of power transmission line engineering data according to claim 1, characterized in that: The construction of the weighted undirected graph further includes: Performing threshold segmentation on the correlation values between any two technical indicators to obtain a first segmentation threshold, and a combination of technical indicators with a correlation value greater than the first segmentation threshold is recorded as having strong correlation; All technical indicators are regarded as nodes, and two technical indicators with strong correlation are connected. The edge weight of the connection is the correlation value of the technical indicator combination. A self-loop is constructed for each technical indicator node, and the edge weight of the self-loop is the aggregation eigenvalue to obtain a weighted undirected graph of all technical indicators.

4. The method for extracting technical indicators of power transmission line engineering data based on AI according to claim 1, characterized in that: The significant screening value of each technical indicator is a normalized result of the product of the node score of each technical indicator in the weighted undirected graph and the degree of the node corresponding to each technical indicator in the weighted undirected graph.

5. The AI-based method for extracting technical indicators of power transmission line engineering data according to claim 1, characterized in that: The classification of significant general technical indicators and non-general technical indicators includes: The significant screening values of all technical indicators are threshold segmented to output a second segmentation threshold, and the technical indicators with significant screening values greater than the second segmentation threshold are regarded as significant general technical indicators, and all other technical indicators are regarded as non-general technical indicators.

6. The AI-based method for extracting technical indicators of power transmission line engineering data according to claim 1, characterized in that: The classification of the transmission line engineering data further includes: The transmission line engineering data are taken as samples, and the normalized results of the numerical values of all significant universal technical indicators of each sample are used to form the universal indicator vector of each sample. A clustering algorithm is used to cluster all transmission line engineering data according to the universal indicator vectors of all samples to obtain clusters.

7. The method for extracting technical indicators of power transmission line engineering data based on AI according to claim 1, characterized in that: The calculation method of the proximity expansion scale of each non-universal technical indicator of each cluster is as follows: Where, is the neighboring expansion scale of the xth non-universal technical indicator of the ith cluster, is the element value of the first mutation point in the inter-cluster distance sequence of the i-th cluster, is the significant screening value of the xth non-general technical indicator, where the inter-class distances between the i-th cluster and all other clusters are arranged in ascending order of values to form the inter-class distance sequence of the i-th cluster.

8. The method for extracting technical indicators of power transmission line engineering data based on AI according to claim 6, characterized in that: Determining the construction contribution weight of each non-general technical indicator of each cluster to obtain a significant scenario-based technical indicator further includes: Extracting adjacent clusters of each cluster through the relationship between the inter-cluster distance and the neighbor expansion scale; Taking the cluster center of any cluster as the target sample, calculate the distance between all samples in the cluster and all its neighboring clusters and the target sample, and sort them in ascending order of distance to extract L samples as the positive sample set; accordingly, calculate the distance between all samples in all clusters except the cluster and all its neighboring clusters and the target sample, and similarly sort them in ascending order of distance to extract L samples as the negative sample set; The calculation method of the construction contribution weight of each non-general technical indicator of each cluster is as follows: Where, The construction contribution weight of the xth non-general technical indicator, 、 They are respectively the positive sample set and the negative sample set. The value of the xth non-general technical indicator for each sample, is the value of the target sample regarding the xth non-general technical indicator; All calculated construction contribution weights are normalized, and non-general technical indicators whose normalized construction contribution weights are greater than or equal to the contribution threshold are regarded as significant scenario-based technical indicators.

9. The method for extracting technical indicators of power transmission line engineering data based on AI according to claim 1, characterized in that: The extracting of adjacent clusters of each cluster further includes: for each cluster, taking clusters with inter-cluster distances less than or equal to the adjacent expansion scale as adjacent clusters of each cluster.

Citation Information

Patent Citations

  • Cluster-analysis-based de-noising method and apparatus for trajectory

    CN106650771A

  • Power transmission and transformation project design document-oriented index extraction optimization method and system

    CN116894176A