Distribution network business expansion project cost budgeting data processing method and system

Through the combination of multidimensional semantic correlation classification, adaptive small sample learning and adversarial generation network, the problems of incomplete information and insufficient abnormal detection in the preparation of cost budget for expansion of the distribution network industry are solved, accurate data classification and dynamic price adjustment are realized, scientific decision-making support is provided, and the accuracy and efficiency of the cost budget is improved.

CN120409882APending Publication Date: 2025-08-01内蒙古电力(集团)有限责任公司内蒙古电力经济技术研究院分公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510310352.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-17
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

The existing distribution network expansion project cost budget preparation technology has problems such as incomplete information acquisition, insufficient data integration, insufficient abnormal detection, untimely price adjustments and insufficient decision-making support, which leads to the disconnection of the budget results from the actual situation and inefficient efficiency.

Method used

Through multi-dimensional semantic correlation classification and adaptive small sample learning classification and sorting data, using adversarial generation network for abnormal detection, combining dynamic adjustment of multi-factor price, defining decision action space and training models, providing scientific decision-making suggestions for engineering, and using off-site storage and incremental backup technologies to ensure data security.

Benefits of technology

It realizes more accurate data classification and abnormal detection, dynamically updates budget data, provides scientific decision-making support, improves the accuracy and efficiency of budget preparation, and reduces resource waste and decision-making errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120409882A_ABST
    Figure CN120409882A_ABST
Patent Text Reader

Abstract

The invention discloses a distribution network business expansion project cost budgeting data processing method and system, and relates to the field of electric power project cost. The method comprises the steps of data collection, data arrangement and classification, data verification, anomaly detection through an adversarial generative network, data adjustment, dynamic adjustment through a multi-factor price, data analysis and application definition decision action space, state information determination, and model training and decision making. The system comprises a memory, a processor and a runnable program, and is also provided with a data backup module based on remote storage and incremental backup technologies. Compared with the prior art, the method can provide comprehensive and accurate budget information, accurate classification data, intelligent verification data and dynamic updating budget, is convenient for auditing analysis, provides scientific decision suggestions for projects, and improves the economic benefits and management level of the projects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of electric power project cost, and specifically to a data processing method and system for preparing the cost budget of distribution network business expansion projects. Background Art

[0002] In the development process of the power system, distribution network business expansion projects play a crucial role. It is a key link to meet the growing power demand of society and improve power supply reliability and service quality. And the preparation of the project cost budget, as one of the core tasks of distribution network business expansion projects, its accuracy and scientificity directly affect the project's cost control, resource allocation and overall economic benefits.

[0003] In the early stage, the preparation of the distribution network business expansion project cost budget mainly relied on manual experience and simple calculation methods. Staff estimated the costs of materials, equipment, labor, etc. required for the project based on their own work experience and limited historical data. This method had many drawbacks. Due to the subjectivity and limitations of experience, the budget results often deviated greatly from the actual costs, easily leading to project cost overruns or resource waste. Moreover, the efficiency of manual calculation was low. When facing complex project structures and a large amount of data, it was not only easy to make calculation errors, but also consumed a lot of time and energy. With the development of computer technology, some simple data processing software began to be applied to the preparation of cost budgets. These software could achieve basic data storage and simple calculation functions, improving the work efficiency and calculation accuracy to a certain extent. However, most of them could only process single-type data and lacked the ability to comprehensively analyze and deeply mine data.

[0004] In terms of data collection, traditional methods have problems of incomplete and untimely information acquisition. The collection of engineering data often relies on limited drawings and documents provided by the design department, and the acquisition of information such as the distribution network line route and pole tower location is not detailed and accurate enough. The collection of market data mainly depends on regular market research and supplier quotations, and the data is not updated in a timely manner, unable to reflect the real-time fluctuations of market prices and supply and demand changes. The collation and utilization of historical engineering data are also not sufficient, lacking systematic analysis and summary, and unable to provide effective reference for current projects. In terms of data classification and collation, existing technical means are relatively simple and crude. Usually, a classification method based on simple attributes is adopted, which is difficult to handle complex data relationships and newly emerging data types. For data from different sources and in different formats, there is a lack of effective integration and unified management, resulting in poor data availability and maintainability. In terms of data verification and adjustment, traditional methods lack an effective anomaly detection mechanism and are difficult to detect errors and outliers in the data. For budget changes caused by factors such as market price fluctuations and engineering changes, they cannot be adjusted in a timely and accurate manner, causing the budget to deviate from the actual situation. In terms of data analysis and application, existing technologies are mainly based on simple statistical analysis and empirical judgment, unable to provide scientific and reasonable decision-making support. For key decision-making links such as construction plans, equipment procurement, and resource allocation, there is a lack of effective models and algorithms for optimization and evaluation, easily leading to decision-making mistakes and resource waste.

[0005] In summary, there are many deficiencies in the existing data processing technologies for the preparation of distribution network service expansion project cost budgets, which cannot meet the needs of the development of modern power engineering. Therefore, there is an urgent need for a more scientific, efficient, and accurate data processing method and system to improve the quality and level of the preparation of distribution network service expansion project cost budgets. Summary of the Invention

[0006] The purpose of the present invention is to provide a data processing method and system for the preparation of distribution network service expansion project cost budgets. By first collecting engineering data, market data, and historical engineering data, then using multi-dimensional semantic association classification and adaptive small-sample learning classification to sort and classify the data, then performing data verification through an adversarial generation network, then adjusting the data with multi-factor price dynamics, and finally realizing data analysis and application by defining the decision action space, determining the state information, model training, and decision-making. At the same time, the system data backup module regularly backs up key data based on off-site storage and incremental backup technologies to solve the above problems.

[0007] To achieve the above purpose, the present invention provides the following technical solutions:

[0008] A data processing method and system for the preparation of distribution network service expansion project cost budgets, including the following steps:

[0009] S1: Data collection; specifically including engineering data collection, obtaining detailed engineering drawings covering the distribution network line routes, pole positions, and cable laying path information from the design department; collecting equipment selection lists to clarify the specifications, models, quantities, and technical parameters of various electrical equipment; obtaining line planning schemes, including the overall architecture, load distribution, and power supply scope of the distribution network;

[0010] Market data collection, establishing connections with major raw material suppliers to obtain real-time data on raw material prices and price fluctuation trends, and collecting market supply and demand data through market research institutions, including the supply quantity, demand quantity, and supply and demand change trends of raw materials;

[0011] Historical engineering data collection, organizing relevant data on previous distribution network business expansion projects within the company, including project costs, construction progress, equipment procurement costs, and human input information.

[0012] S2: Data sorting and classification; by importing engineering drawings, equipment selection lists, and line planning schemes, using multi-dimensional semantic association classification to classify and integrate the data;

[0013] For newly emerging data types, use adaptive small-sample learning classification to dynamically assign reasonable weights to small samples.

[0014] The specific steps of the multi-dimensional semantic association classification are as follows:

[0015] S211: Determination of semantic dimensions and weights; specifically, a team is formed by distribution network engineering experts to jointly determine the semantic dimensions and weights. According to the semantic dimensions covering equipment functional characteristics, technical parameters, production manufacturers and brand reputations, application scenarios, reasonable weights ω k , where k represents the serial number of the semantic dimension, and ω k represents the weight of the k-th semantic dimension;

[0016] S212: Data vector conversion, converting text data into vector representations v i,k and v j,k ; for equipment description texts, use the word vector model Word2Vec to map to a low-dimensional vector space, and directly normalize numerical data and use it as vector elements; at the same time, add timestamps t i and t j , t i represents the time when data i is generated or updated, and t j represents the time when data j is generated or updated;

[0017] S213: Data classification, developing a classification model software based on the Transformer architecture, inputting the data into the model, and calculating the semantic association strength S according to the following formulaij :

[0018]

[0019] Among them, S ij represents the semantic association strength between data i and data j; n is the total number of semantic dimensions; cosinc(v i,k , v j,k ) is the cosine similarity of the vector representations of data i and j in the k-th semantic dimension; ω k is the weight of each dimension, where k represents the serial number of the semantic dimension, ω k represents the weight of the k-th semantic dimension, t i represents the time when data i is generated or updated, t j represents the time when data j is generated or updated, and σ is the time impact coefficient, which is used to measure the influence degree of the time factor on the semantic association strength;

[0020] Set the association strength threshold. When the association strength between the new data and a certain category of data exceeds the threshold, it will be classified into that category.

[0021] The specific steps of the adaptive small-sample learning classification are as follows:

[0022] S221: New data feature extraction. When a new data type appears, identify its key features and performance indicators, and use these indicators as the sample data x l , represents the serial number of the small sample, and x l represents the first small sample data;

[0023] S222: Sample weight calculation. Calculate the sample weight α l according to the following formula:

[0024]

[0025] Among them, m is the number of small samples; ∈ is an extremely small constant used to prevent the denominator from being zero, μ is the overall sample mean, and μ l is the mean of the first small sample; σ l is the standard deviation of the first small sample; is the correlation between the small sample and the overall sample mean;

[0026] S223: Classification training. Adopt the incremental learning algorithm, combine the calculated sample weights, and conduct classification training on the new data to accurately classify it into the appropriate category.

[0027] S3: Data verification; perform anomaly detection through the adversarial generative network, including the following steps:

[0028] S311: Model training. Collect historical power grid connection project data, including normal data and confirmed abnormal data. After uniformly mapping the data values to the interval [0, 1], build a generative adversarial network (GAN) for training.

[0029] S312: Abnormal score calculation. Input the data to be detected into the trained GAN model, and calculate the abnormal score A of the data according to the following formula:

[0030] A = KL(p real ||p gen ) + λ · JS(D(p real )||D(p gen ))

[0031] where KL(p real ||p gen ) is the KL divergence between the real data distribution p real and the generated data distribution p gen , which is used to measure the difference degree between the two distributions; JS(D(p real )||D(p gen )) is the JS divergence of the discriminator D's judgment results on the real data and the generated data, which is also used to measure the difference; λ is the balance coefficient, which is used to balance the roles of the KL divergence and the JS divergence in the abnormal score.

[0032] By analyzing the abnormal score distribution of a large number of normal data, determine the abnormal score threshold. When the abnormal score of the data exceeds the threshold, determine the data as abnormal data, and the system automatically triggers an abnormal alarm to prompt the staff to further analyze the cause of the abnormality.

[0033] S4: Data adjustment; perform dynamic adjustment through multi-factor prices, specifically including the following steps:

[0034] S411: Data tracking and collection. Establish an information sharing mechanism with the main raw material suppliers, and obtain the price change data ΔP m , that is, the raw material price change amount, through electronic data transmission or regular reports; and record the benchmark price P m of the raw materials; collect the market supply and demand change amounts ΔS and ΔD using market research reports, which are the market supply change amount and the market demand change amount respectively; and the benchmark supply and demand quantities S and D, that is, the market benchmark supply quantity and the benchmark demand quantity.

[0035] S412: Price adjustment calculation. Substitute the collected data into the following formula to calculate the price adjustment coefficient γ:

[0036]

[0037] where α is the raw material price impact weight, β and δ are the supply and demand impact weights respectively, and ΔPm is the raw material price, P m is the benchmark price of raw materials, ΔS and ΔD are the changes in market supply and demand, and S and D are the benchmark supply and demand volumes;

[0038] Apply the adjustment coefficient to the equipment material prices in the budget data to achieve dynamic price updates, record the relevant information of each price adjustment, including the adjustment time, reason, and prices before and after the adjustment, and form a price adjustment record file for subsequent auditing and cost analysis.

[0039] S5: Data analysis and application. Specifically, it includes the following steps:

[0040] S511: Define the decision action space A, including construction plan decisions such as different construction techniques, construction sequences, and equipment installation methods; equipment procurement decisions such as procurement time, supplier selection, and procurement quantity; resource allocation strategies including human, material, and financial resources;

[0041] S512: Determine the state information S t , including project progress, cost data, and market conditions. t represents the time step number, and S t represents the state information at time step t;

[0042] S513: Model training: Use historical project data and simulation data to train the model. During the training process, calculate the value V(A) of the decision action A according to the following formula:

[0043]

[0044] where T is the total number of decision time steps; λ is the discount factor, used to measure the importance of future rewards; R t (A) is the immediate reward obtained by taking the decision action A at time step t; u is the total number of deep neural networks; ∈ k is the weight coefficient adjusted according to network importance and accuracy, and k represents the serial number of the deep neural network; F k (A,S t ) is the value evaluation function output by the kth deep neural network according to the decision action A and state S t ; Entropy(A) is the entropy of the decision action A, reflecting the uncertainty of the decision; Var(R) is the variance of the reward, used to measure the fluctuation degree of the reward; σ is the variance influence coefficient; By adjusting the weights of the neural network and the reward function of reinforcement learning, enable the agent to learn the optimal decision-making strategy;

[0045] S514: Model decision-making. In actual projects, according to the current state information S t , use the trained model to provide decision-making suggestions for the project.

[0046] A data processing system for preparing a construction cost budget for network expansion in the distribution network, characterized in that it includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, it implements the steps of the method for preparing a construction cost budget for network expansion in the distribution network according to any one of claims 1-8 above.

[0047] The system includes a data backup module that regularly backs up critical data in the system based on off-site storage and incremental backup technologies.

[0048] The method and system for preparing a construction cost budget for network expansion in the distribution network first start from three aspects: engineering data, market data, and historical engineering data in the data collection stage to obtain all-round information related to the project, laying a foundation for subsequent processing. Then it enters the data sorting and classification link. Using multi-dimensional semantic association classification, experts determine the semantic dimensions and weights, convert the data into vector form and consider time factors, and calculate the semantic association strength through a model based on the Transformer architecture for classification; for newly emerging data types, adaptive small-sample learning classification is used, features are extracted, sample weights are calculated, and then an incremental learning algorithm is used for training and classification. In the subsequent data verification, an adversarial generative network is used. The model is trained to learn the characteristics of normal and abnormal data, and the anomaly score of the data to be detected is calculated. If it exceeds the threshold, it is determined as abnormal and an alarm is triggered. In the data adjustment stage, an information sharing mechanism is established to track and collect price and supply-demand data, calculate the price adjustment coefficient to update the budget data and record relevant information. Finally, in data analysis and application, the decision action space is defined and the state information is determined. The model is trained using historical and simulated data, the decision action value is calculated, and in actual projects, the trained model provides decision-making suggestions based on the state information. And the data backup module in the system regularly backs up critical data based on off-site storage and incremental backup technologies to ensure the data security and recoverability of the entire process.

[0049] Compared with the prior art, the beneficial effects of the present invention are:

[0050] It comprehensively covers engineering data, market data, and historical engineering data, and at the same time sorts out the historical engineering data within the company, which can provide more comprehensive and up-to-date information for budget preparation, making the budget more in line with the actual situation. And it fully considers the internal relationships between various data, such as the association between engineering data and market data, historical engineering data, and can more accurately grasp the influencing factors of project costs, providing a more targeted data basis for subsequent processing.

[0051] By classifying through multi-dimensional semantic association, the semantic dimensions and weights are determined by distribution network engineering experts, and the semantic association strength is calculated in combination with the Transformer architecture model for classification, which can more accurately reflect the semantic relationship between data and improve the classification accuracy. For newly emerging data types, the adaptive small-sample learning classification can dynamically assign reasonable weights to small samples, effectively handle the classification problem of small-sample data, and avoid inaccurate classification caused by data type changes.

[0052] This method uses a generative adversarial network for anomaly detection. By training the model to learn the characteristics of normal and abnormal data and calculating the anomaly score of the data to be detected, it can automatically adapt to data changes, achieve intelligent anomaly detection, reduce manual intervention, and improve the verification efficiency. It can more accurately discover anomalies in the data, trigger alarms in a timely manner, and provide more reliable guarantees for data quality.

[0053] By establishing an information sharing mechanism with suppliers, tracking the changes in raw material prices and supply and demand in real time, and using a multi-factor price formula to calculate the adjustment coefficient, the dynamic update of budget data can be realized, so that the budget can better reflect the actual market situation and effectively control the project cost.

[0054] Record detailed information during each price adjustment to form a price adjustment record file, which is convenient for subsequent audits and cost analysis. Compared with the existing technologies lacking data recording and analysis mechanisms, it can better summarize experiences and lessons and provide more valuable references for future projects.

[0055] This method defines the decision action space, determines the state information, trains the model using historical and simulation data, and comprehensively considers various factors to calculate the value of decision actions, which can provide more scientific and reasonable decision-making suggestions for construction plans, equipment procurement, resource allocation, etc., and improve the economic benefits and management level of the project. Brief Description of the Drawings

[0056] Figure 1 It is a flowchart of a data processing method for preparing the project cost budget of distribution network service expansion in the present invention. Specific Implementation Method

[0058] Next, the technical solutions in the embodiments of the present invention will be completely described in conjunction with the accompanying drawings in the embodiments of the present invention.

[0059] As Figure 1 shown, a data processing method for preparing the project cost budget of distribution network service expansion is characterized by including the following steps:

[0060] S1: Data collection;

[0061] S2: Data sorting and classification;

[0062] S3: Data verification;

[0063] S4: Data adjustment;

[0064] S5: Data analysis and application.

[0065] The specific steps of data collection in step S1 include the following parts:

[0066] Collection of engineering data. For the collection of engineering data, obtain detailed engineering drawings covering the distribution network line route, pole positions, and cable laying path information from the design department; collect the equipment selection list to clarify the specifications, models, quantities, and technical parameters of various electrical equipment; obtain the line planning scheme, including the overall architecture, load distribution, and power supply scope of the distribution network;

[0067] Collection of market data. Establish contact with major raw material suppliers to obtain real-time data on raw material prices and price fluctuation trends, and collect market supply and demand data through market research institutions, including the supply quantity, demand quantity, and supply and demand change trends of raw materials;

[0068] Collection of historical engineering data. Sort out the relevant data of past distribution network business expansion projects within the company, including project cost, construction progress, equipment procurement cost, and manpower input information.

[0069] In step S2, data sorting is carried out by importing engineering drawings, equipment selection lists, and line planning schemes, and the data is classified and integrated using multi-dimensional semantic association classification;

[0070] For newly emerging data types, adaptive small-sample learning classification is used to dynamically assign reasonable weights to small samples.

[0071] The specific steps of the multi-dimensional semantic association classification are as follows:

[0072] S211: Determination of semantic dimensions and weights. Specifically, a team is formed by distribution network engineering experts to jointly determine the semantic dimensions and weights. According to the semantic dimensions covering equipment functional characteristics, technical parameters, manufacturer and brand reputation, and application scenarios, reasonable weights ω k , where k represents the serial number of the semantic dimension, and ω k represents the weight of the kth semantic dimension;

[0073] S212: Conversion of data vectors. Convert the text data into vector representations v i,k and v j,k ; for equipment description texts, use the word vector model Word2Vec to map them to a low-dimensional vector space, and directly normalize the numerical data and use it as vector elements; at the same time, add time stamps t i and t j , t iIndicates the time when data i is generated or updated, t j Indicates the time when data j is generated or updated;

[0074] S213: Data classification. Develop a classification model software based on the Transformer architecture, input the data into the model, and calculate the semantic association strength S according to the following formula ij :

[0075]

[0076] Among them, S ij Indicates the semantic association strength between data i and data j; n is the total number of semantic dimensions; cosinc(v i,k , v j,k ) is the cosine similarity of the vector representations of data i and j in the k-th semantic dimension; ω k Is the weight of each dimension, where k represents the serial number of the semantic dimension, ω k Represents the weight of the k-th semantic dimension, t i Indicates the time when data i is generated or updated, t j Indicates the time when data j is generated or updated, and σ is the time impact coefficient, which is used to measure the impact degree of time factors on the semantic association strength;

[0077] Set the association strength threshold. When the association strength between new data and a certain category of data exceeds the threshold, classify it into that category.

[0078] The specific steps of the adaptive small-sample learning classification are as follows:

[0079] S221: New data feature extraction. When a new data type appears, identify its key features and performance indicators, and use these indicators as sample data x l , Indicates the serial number of the small sample, x l Represents the first small sample data;

[0080] S222: Sample weight calculation. Calculate the sample weight α according to the following formula l :

[0081] I

[0082] Among them, m is the number of small samples; ∈ is a very small constant used to prevent the denominator from being zero, μ is the overall sample mean, μ l Is the mean of the first small sample; σ l Is the standard deviation of the first small sample; Is the correlation between the small sample and the overall sample mean;

[0083] S223: Classification training. An incremental learning algorithm is adopted, and combined with the calculated sample weights, new data is classified and trained to accurately classify it into the appropriate category.

[0084] In step S3, data verification is to perform anomaly detection through an adversarial generation network, including the following steps:

[0085] S311: Model training. Historical power distribution project data is collected, including normal data and confirmed abnormal data. After uniformly mapping the data values to the interval [0, 1], a generative adversarial network (GAN) is built for training;

[0086] S312: Anomaly score calculation. The data to be detected is input into the trained GAN model, and the anomaly score A of the data is calculated according to the following formula:

[0087] A = KL(p real ||p gen ) + λ · JS(D(p real )||D(p gen ))

[0088] where KL(p real ||p gen ) is the KL divergence between the true data distribution p real and the generated data distribution p gen , which is used to measure the difference degree between the two distributions; JS(D(p real )||D(p gen )) is the JS divergence of the discriminator D's judgment results on the true data and the generated data, which is also used to measure the difference; λ is a balance coefficient, which is used to balance the roles of the KL divergence and the JS divergence in the anomaly score;

[0089] By analyzing the anomaly score distribution of a large number of normal data, the anomaly score threshold is determined; when the anomaly score of the data exceeds the threshold, the data is determined to be abnormal data, and the system automatically triggers an anomaly alarm to prompt the staff to further analyze the cause of the anomaly.

[0090] [[ID=4`1]]In step S4, data adjustment is to perform dynamic adjustment through multi-factor prices, specifically including the following steps:

[0091] S411: Data tracking and collection. An information sharing mechanism is established with the main raw material suppliers, and the price change data ΔP m , that is, the raw material price change amount, is obtained through electronic data transmission or regular reports; and the benchmark price P m of the raw material is recorded; the market supply and demand change amounts ΔS and ΔD are collected using market research reports, which are the market supply change amount and the market demand change amount respectively; and the benchmark supply and demand quantities S and D, that is, the market benchmark supply quantity and the benchmark demand quantity;

[0092] S412: Price adjustment calculation. Substitute the collected data into the following formula to calculate the price adjustment coefficient γ:

[0093]

[0094] where α is the weight of raw material price impact, β and δ are the weights of supply and demand impact respectively, ΔP m is the current raw material price, P m is the benchmark price of the raw material, ΔS and ΔD are the changes in market supply and demand, and S and D are the benchmark supply and demand volumes;

[0095] Apply the adjustment coefficient to the equipment material prices in the budget data to achieve dynamic price updates. Record the relevant information of each price adjustment, including the adjustment time, reason, and prices before and after the adjustment, to form a price adjustment record file for subsequent auditing and cost analysis.

[0096] Step S5: Data analysis and application specifically includes the following steps:

[0097] S511: Define the decision action space A, including construction plan decisions such as different construction techniques, construction sequences, and equipment installation methods; equipment procurement decisions including procurement time, supplier selection, and procurement quantity; resource allocation strategies including human, material, and financial resources;

[0098] S512: Determine the state information S t , including project progress, cost data, and market conditions. t represents the time step sequence number, and S t represents the state information at time step t;

[0099] S513: Model training: Use historical project data and simulation data to train the model. During the training process, calculate the value V(A) of the decision action A according to the following formula:

[0100]

[0101] where T is the total number of decision time steps; λ is the discount factor, used to measure the importance of future rewards; R t (A) is the immediate reward obtained by taking the decision action A at time step t; u is the total number of deep neural networks; ∈ k is the weight coefficient adjusted according to network importance and accuracy, and k represents the serial number of the deep neural network; F k (A,S t ) is the kth deep neural network based on the decision action A and state S tOutput value evaluation function; Entropy(A) is the entropy of decision-making action A, reflecting the uncertainty of decision-making; Var(R) is the variance of the reward, used to measure the degree of reward fluctuation; σ is the variance influence coefficient; by adjusting the weights of the neural network and the reward function of reinforcement learning, the agent learns the optimal decision-making strategy;

[0102] S514: Model decision-making. In actual engineering, based on the current state information S t , use the trained model to provide decision-making suggestions for the project.

[0103] A data processing system for preparing the engineering cost budget of distribution network service expansion. It is characterized by including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, it realizes the steps of the data processing method for preparing the engineering cost budget of distribution network service expansion according to any one of claims 1-8 above.

[0104] The system includes a data backup module, which regularly backs up the key data in the system based on off-site storage and incremental backup technologies.

[0105] The specific implementation manners of the data processing method and system for preparing the engineering cost budget of distribution network service expansion are as follows:

[0106] Firstly, it is the data collection stage. In terms of collecting engineering materials, obtain detailed engineering drawings containing information such as the distribution network line route, pole position, and cable laying path from the design department, collect an equipment selection list that clarifies the specifications, models, quantities, and technical parameters of various electrical equipment, and a line planning scheme covering the overall distribution network architecture, load distribution, and power supply scope; for market data collection, establish contacts with major raw material suppliers to obtain real-time prices and fluctuation trends, and collect supply and demand data of raw materials through market research institutions; for historical engineering data collection, sort out information such as the engineering cost, construction progress, equipment procurement cost, and labor input of previous distribution network service expansion projects within the company.

[0107] Then enter data sorting and classification. By importing engineering drawings, equipment selection lists, and line planning schemes, use multi-dimensional semantic association classification to classify and integrate the data;

[0108] For newly emerging data types, use adaptive small-sample learning classification to dynamically assign reasonable weights to small samples.

[0109] The specific steps of the multi-dimensional semantic association classification are as follows:

[0110] S211: Determine the semantic dimension and weight. Specifically, a team is formed by distribution network engineering experts to jointly determine the semantic dimension and weight. According to the equipment functional characteristics, technical parameters, manufacturers and brand reputation, and application scenarios covered by the semantic dimension, reasonable weights ω are assigned to each dimension according to the actual engineering requirements and experience, where k represents the serial number of the semantic dimension, and ω represents the weight of the k-th semantic dimension. k , where k represents the serial number of the semantic dimension, and ω k represents the weight of the k-th semantic dimension;

[0111] S212: Convert data vectors. Convert the text data into vector representations v and v; for the equipment description text, use the word vector model Word2Vec to map it into a low-dimensional vector space, and directly normalize the numerical data and use it as vector elements; at the same time, add time stamps t and t to each data record, where t represents the time when data i is generated or updated, and t represents the time when data j is generated or updated. i,k and v j,k ; for the equipment description text, use the word vector model Word2Vec to map it into a low-dimensional vector space, and directly normalize the numerical data and use it as vector elements; at the same time, add time stamps t i and t j , t i represents the time when data i is generated or updated, and t j represents the time when data j is generated or updated;

[0112] S213: Data classification. Develop a classification model software based on the Transformer architecture, input the data into the model, and calculate the semantic association strength S according to the following formula: ij :

[0113]

[0114] where, S ij represents the semantic association strength between data i and data j; n is the total number of semantic dimensions; cosinc(v i,k , v j,k ) is the cosine similarity of the vector representations of data i and j in the k-th semantic dimension; ω k is the weight of each dimension, where k represents the serial number of the semantic dimension, and ω k represents the weight of the k-th semantic dimension, t i represents the time when data i is generated or updated, and t j represents the time when data j is generated or updated, and σ is the time impact coefficient, which is used to measure the impact degree of time factors on the semantic association strength;

[0115] Set the association strength threshold. When the association strength between the new data and a certain category of data exceeds the threshold, classify it into this category.

[0116] The specific steps of the adaptive small-sample learning classification are as follows:

[0117] S221: Extract new data features. When a new data type appears, identify its key features and performance indicators, and use these indicators as sample data xl , represents the serial number of a small sample, x l represents the data of the first small sample;

[0118] S222: Sample weight calculation. Calculate the sample weight α according to the following formula l :

[0119]

[0120] where m is the number of small samples; ∈ is an extremely small constant used to prevent the denominator from being zero, μ is the overall sample mean, μ l is the mean of the first small sample; σ l is the standard deviation of the first small sample; is the correlation between the small sample and the overall sample mean;

[0121] S223: Classification training. Adopt an incremental learning algorithm, combine the calculated sample weights, and conduct classification training on new data to accurately classify it into the appropriate category.

[0122] In the data verification link, through the generative adversarial network, collect historical normal and abnormal data, map them to the [0,1] interval to train the GAN model, input the data to be detected into the model to calculate the anomaly score, and if it exceeds the threshold, it is determined as abnormal and an alarm is triggered. It includes the following steps:

[0123] S311: Model training. Collect historical power distribution network project data, including normal data and confirmed abnormal data, map the data values to the [0,1] interval uniformly, and then build a generative adversarial network (GAN) for training;

[0124] S312: Anomaly score calculation. Input the data to be detected into the trained GAN model, and calculate the anomaly score A of the data according to the following formula:

[0125] A = KL(p real ||p gen ) + λ·JS(D(p real )||D(p gen ))

[0126] where KL(p real ||p gen ) is the KL divergence between the true data distribution p real and the generated data distribution p gen , which is used to measure the difference degree between the two distributions; JS(D(p real )||D(p gen)) is the Jensen-Shannon divergence of the discriminator D's judgment results on real data and generated data, which is also used to measure the difference; λ is the balance coefficient, which is used to balance the roles of KL divergence and JS divergence in anomaly scoring;

[0127] By analyzing the anomaly score distribution of a large number of normal data, the anomaly score threshold is determined; when the anomaly score of the data exceeds the threshold, the data is determined to be abnormal data, and the system automatically triggers an anomaly alarm to prompt the staff to further analyze the cause of the anomaly.

[0128] Data adjustment is carried out through multi-factor price dynamic adjustment, establishing an information sharing mechanism with suppliers to collect price changes and supply-demand change data, calculating the price adjustment coefficient and applying it to the budget data, and at the same time recording the adjustment information to form a file. The specific steps are as follows:

[0129] S411: Data tracking and collection, establishing an information sharing mechanism with the main raw material suppliers, and obtaining the price change data ΔP m , that is, the raw material price change amount; and recording the benchmark price P of the raw material m ; using market research reports to collect the market supply-demand change amounts ΔS and ΔD, which are the market supply change amount and market demand change amount respectively; and the benchmark supply-demand quantities S and D, that is, the market benchmark supply quantity and benchmark demand quantity;

[0130] S412: Price adjustment calculation, substituting the collected data into the following formula to calculate the price adjustment coefficient γ:

[0131]

[0132] where α is the weight of the raw material price impact, β and δ are the supply-demand impact weights respectively, ΔP m is the raw material price, P m is the benchmark price of the raw material, ΔS and ΔD are the market supply-demand change amounts, and S and D are the benchmark supply-demand quantities;

[0133] Apply the adjustment coefficient to the equipment material price in the budget data to achieve dynamic price update, record the relevant information of each price adjustment, including the adjustment time, adjustment reason, and prices before and after adjustment, to form a price adjustment record file for subsequent auditing and cost analysis.

[0134] Finally, in the data analysis and application stage, first define the decision-making action space such as construction plans, equipment procurement, and resource allocation, determine the state information such as project progress, cost data, and market conditions, use historical and simulated data to train the model and calculate the value of decision-making actions, and enable the intelligent agent to learn the optimal strategy by adjusting the neural network weights and reinforcement learning reward functions. In actual projects, provide decision-making suggestions based on the current state information by the trained model. S511: Define the decision-making action space A, including construction plan decisions including different construction techniques, construction sequences, and equipment installation methods; equipment procurement decisions including procurement time, supplier selection, and procurement quantity; resource allocation strategies including human, material, and financial resources;

[0135] S512: Determine the state information S t , including project progress, cost data, and market conditions. t represents the time step number, and S t represents the state information at time step t;

[0136] S513: Model training: Use historical project data and simulated data to train the model. During the training process, calculate the value V(A) of the decision-making action A according to the following formula:

[0137]

[0138] where T is the total number of decision time steps; λ is the discount factor, used to measure the importance of future rewards; R t (A) is the immediate reward obtained by taking the decision-making action A at time step t; u is the total number of deep neural networks; ∈ k is the weight coefficient adjusted according to network importance and accuracy, and k represents the serial number of the deep neural network; F k (A,S t ) is the value evaluation function output by the kth deep neural network according to the decision-making action A and the state S t ; Entropy(A) is the entropy of the decision-making action A, reflecting the uncertainty of the decision; Var(R) is the variance of the reward, used to measure the fluctuation degree of the reward; σ is the variance influence coefficient; By adjusting the weights of the neural network and the reward function of reinforcement learning, enable the intelligent agent to learn the optimal decision-making strategy;

[0139] S514: Model decision-making. In actual projects, based on the current state information S t , use the trained model to provide decision-making suggestions for the project.

Claims

1. A data processing method for preparing the engineering cost budget of distribution network business expansion, characterized in that, It includes the following steps: S1: Data collection; S2: Data sorting and classification; S3: Data verification; S4: Data adjustment; S5: Data analysis and application.

2. A data processing method for preparing the engineering cost budget of distribution network business expansion according to claim 1, characterized in that, The data collection in step S1 specifically includes the following parts: Engineering data collection. For engineering data collection, obtain detailed engineering drawings covering the distribution network line route, pole position, and cable laying path information from the design department; collect the equipment selection list to clarify the specifications, models, quantities, and technical parameters of various electrical equipment; obtain the line planning scheme, including the overall architecture, load distribution, and power supply scope of the distribution network; Market data collection. Establish contact with major raw material suppliers to obtain real-time data on raw material prices and price fluctuation trends, and collect market supply and demand data through market research institutions, including the supply quantity, demand quantity, and supply and demand change trends of raw materials; Historical engineering data collection. Sort out the relevant data of previous distribution network business expansion projects within the company, including project cost, construction progress, equipment procurement cost, and manpower input information.

3. A data processing method for preparing the engineering cost budget of distribution network business expansion according to claim 1, characterized in that In step S2, data sorting classifies and integrates the data through multi-dimensional semantic association classification by importing engineering drawings, equipment selection lists, and line planning schemes; For newly emerging data types, use adaptive small-sample learning classification to dynamically assign reasonable weights to small samples.

4. A data processing method for preparing a project cost budget for network expansion in power distribution, as described in claim 3, characterized in that The specific steps of the multi-dimensional semantic association classification are as follows: S211: Determine the semantic dimension and weight. Specifically, a team is formed by distribution network engineering experts to jointly determine the semantic dimension and weight. According to the equipment functional characteristics, technical parameters, manufacturer and brand reputation, and application scenarios covered by the semantic dimension, reasonable weights ω are assigned to each dimension according to the actual engineering requirements and experience. k , where k represents the serial number of the semantic dimension, and ω k represents the weight of the k-th semantic dimension; S212: Data vector transformation, transforming text data into vector representation v i,k and v j,k ; For device description text, it is mapped to a low-dimensional vector space using the word vector model Word2Vec, and numerical data is directly normalized and used as vector elements; at the same time, a timestamp t i and t j , t i represents the time when data i is generated or updated, and t j represents the time when data j is generated or updated; S213: Data classification. Develop a classification model software based on the Transformer architecture, input the data into the model, and calculate the semantic association strength S according to the following formula ij : Among them, S ij represents the semantic association strength between data i and data j; n is the total number of semantic dimensions; cosinc(v i,k , v j,k ) is the cosine similarity of the vector representations of data i and j in the k-th semantic dimension; ω k is the weight of each dimension, where k represents the serial number of the semantic dimension, ω k represents the weight of the k-th semantic dimension, t i represents the time when data i is generated or updated, t j represents the time when data j is generated or updated, and σ is the time impact coefficient, which is used to measure the influence degree of time factors on the semantic association strength; Set the association strength threshold. When the association strength between new data and a certain category of data exceeds the threshold, classify it into this category.

5. A data processing method for preparing a project cost budget for network expansion in distribution networks according to claim 3, characterized in that, The specific steps of the adaptive small-sample learning classification are as follows: S221: New data feature extraction. When a new data type appears, identify its key features and performance indicators, and use these indicators as sample data x l , represents the serial number of a small sample, and x l represents the first small sample data; S222: Sample weight calculation, calculate the sample weight α according to the following formula l :[[]]END]] Among them, m is the number of small samples; ∈ is an extremely small constant used to prevent the denominator from being zero, μ is the overall sample mean, μ l is the mean of the first small sample; σ l is the standard deviation of the first small sample; is the correlation between the small sample and the overall sample mean; S223: Classification training. Use the incremental learning algorithm and combine the calculated sample weights to perform classification training on new data and accurately classify it into the appropriate category.

6. A data processing method for preparing the engineering cost budget of the distribution network business expansion according to claim 1, characterized in that, In step S3, data verification performs anomaly detection through a generative adversarial network, including the following steps: S311: Model training. Collect historical distribution network project data, including normal data and confirmed abnormal data. After uniformly mapping the data values to the [0, 1] interval, build a generative adversarial network (GAN) for training; S312: Anomaly score calculation. Input the data to be detected into the trained GAN model, and calculate the anomaly score A of the data according to the following formula: A = KL(p real || p gen ) + λ·JS(D(p real ) || D(p gen )) where KL(p real || p gen ) is the KL divergence between the true data distribution p real and the generated data distribution p gen , which is used to measure the difference degree between the two distributions; JS(D(p real )) ||| D(p gen )) is the JS divergence of the discriminator D's judgment results on the true data and the generated data, which is also used to measure the difference; λ is the balance coefficient, which is used to balance the roles of the KL divergence and the JS divergence in the anomaly score; By analyzing the anomaly score distribution of a large number of normal data, determine the anomaly score threshold; when the anomaly score of the data exceeds the threshold, determine the data as abnormal data, and the system automatically triggers an anomaly alarm to prompt the staff to further analyze the cause of the anomaly.

7. A data processing method for preparing a project cost budget for network expansion in distribution networks according to claim 1, characterized in that In step S4, data adjustment is dynamically adjusted through multi-factor prices, specifically including the following steps: S411: Data tracking and collection, establishing an information sharing mechanism with major raw material suppliers, obtaining price change data ΔP through electronic data transmission or regular reports m , that is, the raw material price change amount; and recording the benchmark price P of the raw material m ; Using market research reports to collect market supply and demand change amounts ΔS and ΔD, which are the market supply change amount and market demand change amount respectively; and the benchmark supply and demand quantities S and D, that is, the market benchmark supply quantity and benchmark demand quantity; S412: Price adjustment calculation. Substitute the collected data into the following formula to calculate the price adjustment coefficient γ: Among them, α is the weight of the impact of raw material prices, β and δ are the weights of supply and demand impacts respectively, and ΔP m is the raw material price, P m is the benchmark price of the raw material, ΔS and ΔD are the changes in market supply and demand, and S and D are the benchmark supply and demand volumes; Apply the adjustment coefficient to the equipment material prices in the budget data to achieve dynamic price update, record the relevant information of each price adjustment, including the adjustment time, adjustment reason, and prices before and after adjustment, and form a price adjustment record file for subsequent auditing and cost analysis.

8. A data processing method for preparing the engineering cost budget of distribution network business expansion according to claim 1, characterized in that The data analysis and application in step S5 specifically include the following steps: S511: Define the decision-making action space A, including construction plan decisions, such as different construction techniques, construction sequences, and equipment installation methods; equipment procurement decisions, including procurement time, supplier selection, and procurement quantity; resource allocation strategies, including human resources, material resources, and financial resources. S512: Determine the status information S t , including the project progress, cost data, and market conditions. t represents the time step number, and S t represents the status information at time step t; S513: Model training: Use historical engineering data and simulation data to train the model. During the training process, calculate the value V(A) of the decision-making action A according to the following formula: Among them, T is the total number of decision time steps; λ is the discount factor, which is used to measure the importance of future rewards; R t (A) is the immediate reward obtained by taking the decision action A at time step t; u is the total number of deep neural networks; ∈ k is the weight coefficient adjusted according to network importance and accuracy, k represents the serial number of the deep neural network; F k (A,S t ) is the value evaluation function output by the k-th deep neural network according to the decision action A and the state S t ; Entropy(A) is the entropy of the decision action A, which reflects the uncertainty of the decision; Var(R) is the variance of the reward, which is used to measure the degree of reward fluctuation; σ is the variance influence coefficient; by adjusting the weights of the neural network and the reward function of reinforcement learning, the agent can learn the optimal decision-making strategy; S514: Model decision-making. In actual engineering, based on the current status information S t , use the trained model to provide decision-making suggestions for the project.

9. A data processing system for preparing the engineering cost budget of distribution network business expansion, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is loaded into the processor, it realizes the steps of the data processing method for compiling the engineering cost budget of the distribution network service expansion according to any one of claims 1-8 above.

10. A data processing system for preparing a project cost budget for network expansion in power distribution, as claimed in claim 9, wherein, The system includes a data backup module, which regularly backs up the key data in the system based on off-site storage and incremental backup technologies.