A method, storage medium and system for identifying key industrial chain nodes in a park based on multi-source data

By processing and clustering multi-source data from enterprises in the park, key industrial chain nodes within the park are identified, solving the problem of enterprise selection and introduction during the park's investment promotion process, improving the adaptability of the industrial chain within the park, and ensuring stable and orderly development.

CN116842407BActive Publication Date: 2026-01-20CHINA SOUTHERN POWER GRID BIG DATA SERVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310799437.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-30
Publication Date
2026-01-20
Estimated Expiration
2043-06-30

AI Technical Summary

Technical Problem

At present, there is a lack of effective means to identify the key industrial chain nodes that need to be introduced into the park, resulting in a low degree of fit between the park's investment promotion goals and the industrial chain, making it difficult to achieve stable and orderly long-term development.

Method used

By acquiring multi-source data from various enterprises within the park, including a list of park enterprises, basic enterprise information data, and power contract data, the data is fused and aligned. Complex network models and semantic recognition technologies are used to analyze the nodes in the industrial chain, calculate multiple key identification indicators, and perform cluster analysis to identify key nodes in the industrial chain.

Benefits of technology

Identifying key industrial chain nodes that need to be introduced into the park provides a reference for the selection and introduction of enterprises during the park's investment promotion process, improves the fit between the park's investment promotion goals and the industrial chain, and achieves stable and orderly long-term development.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116842407B_ABST
    Figure CN116842407B_ABST
Patent Text Reader

Abstract

This invention provides a method, storage medium, and system for identifying key industrial chain nodes within a park based on multi-source data. The method includes: acquiring multi-source data from various enterprises within the park from different data systems; performing fusion and alignment processing on the multi-source data of each enterprise; analyzing the industrial chain to which each enterprise belongs and its node within that chain based on the multi-source data; calculating multiple key identification indicators for each enterprise based on the multi-source data; and performing cluster analysis and key industrial chain node identification based on these key identification indicators. This allows for the identification of key industrial chain nodes within the park that require enterprise introduction, providing a reference for enterprise selection and introduction during the park's investment promotion process, improving the fit between the park's investment promotion goals and the industrial chain within the park, and ultimately achieving stable and orderly long-term development of the park.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, in particular to a method for identifying key industrial chain nodes in a park based on multi-source data, a storage medium and a system. BACKGROUND

[0002] Industrial clusters are a phenomenon in which enterprises are gathered together according to the connection of industrial chains in a region dominated by a central product. How to develop industrial clusters to improve the competitiveness of regional economy is a strategic issue in economic development. Parks are an important part of developing industrial clusters, and enterprises gather in the form of industrial chains to improve the competitiveness of regional economy. Park investment is an important way to develop parks, and plays an important role in the industrial agglomeration and the final development of the park. However, there is currently a lack of reference means for screening and introducing enterprises in the process of park investment, and it is difficult to identify key industrial chain nodes in the park. This leads to a low degree of adaptation between the park investment target and the industrial chain in the park, and makes it difficult to achieve stable and orderly long-term development of the park. SUMMARY

[0003] The technical problem to be solved by the present application is how to identify key industrial chain nodes in the park.

[0004] To solve the above technical problems, the present application provides a method for identifying key industrial chain nodes in a park based on multi-source data, comprising the following steps:

[0005] S1. Obtain multi-source data of each enterprise in the park from different data systems, and perform fusion alignment processing on the multi-source data of each enterprise in the park, wherein the multi-source data includes a park enterprise list, enterprise basic information data and power contract data;

[0006] S2. Analyze the industrial chain to which each enterprise in the park belongs and the industrial chain node in the industrial chain according to the multi-source data;

[0007] S3. Calculate a plurality of key identification indexes of each enterprise in the park according to the multi-source data, wherein the key identification indexes include the competitiveness level of each enterprise in the park, the monthly average power consumption capacity value, the power consumption industry expansion demand score and the daily average power consumption capacity value of the new park enterprise, and specifically include the following steps S31, S32, S33 and S34;

[0008] S31. Based on the complex network model, network nodes corresponding to each enterprise in the park are generated according to the enterprise list in the park, then a similar enterprise network is constructed according to the similarity between the network nodes and each enterprise in the park, and the competition level of each enterprise in the park is analyzed according to the similar enterprise network, wherein the similarity between each enterprise in the park is specifically calculated according to the enterprise basic information data of each enterprise in the park;

[0009] S32. The theoretical monthly average power consumption of each enterprise in the park is calculated according to the power contract data of each enterprise in the park, the actual monthly average power consumption of each enterprise in the park is obtained, and the actual monthly average power consumption of each enterprise in the park is compared with the theoretical monthly average power consumption to obtain the monthly average power consumption overcapacity value of each enterprise in the park;

[0010] S33. According to the power consumption or annual maximum load of each enterprise in the park within a certain time range, the predicted power consumption and the predicted annual maximum load of each enterprise in the park within a future preset time are predicted, and the prediction results are compared with the current annual power consumption and the current annual maximum load, respectively, and the comparison results are analyzed to obtain the power expansion demand score of each enterprise in the park;

[0011] S34. According to the enterprise list in the park and the power contract data of each enterprise in the park, a new enterprise entering the park is identified, and the theoretical daily average power consumption of the new enterprise entering the park is calculated according to the power contract data, the actual daily average power consumption of the new enterprise entering the park is obtained, and the actual daily average power consumption of the new enterprise entering the park is compared with the theoretical daily average power consumption to obtain the daily average power consumption overcapacity value of the new enterprise entering the park;

[0012] S4. According to each key identification index of each enterprise in the park, clustering analysis and key industrial chain node identification of each industrial chain node in the park are performed, specifically including the following steps S41, S42 and S43;

[0013] S41. The average value of each key identification index of each industrial chain node is calculated;

[0014] S42. The clustering analysis of the industrial chain nodes to which each enterprise in the park belongs is performed according to the average value of each key identification index of each enterprise in the park, so as to divide all the industrial chain nodes into a preset number of different categories;

[0015] S43. The average value of each key identification index of all industrial chain nodes in each category after clustering analysis is obtained, and each category is identified as a key industrial chain node according to the obtained result, and a corresponding visual label is generated.

[0016] Preferably, the step S42 comprises:

[0017] S421. According to the business needs, initially determine the number of categories to be clustered for all industry chain nodes;

[0018] S422. Randomly define the cluster centers of each category on different industry chain nodes;

[0019] S423. Calculate the continuous type classification factor distance from the non-cluster center industry chain nodes to each cluster center;

[0020] S424. Calculate the discrete type classification factor distance from the non-cluster center industry chain nodes to each cluster center;

[0021] S425. Calculate the distance from the non-cluster center industry chain nodes to each cluster center according to the continuous type classification factor distance and the discrete type classification factor distance;

[0022] S426. Find the minimum value among the distances from each non-cluster center industry chain node to each cluster center, and then associate each non-cluster center industry chain node to the category to which the nearest cluster center belongs;

[0023] S427. After associating all non-cluster center industry chain nodes to the categories to which the nearest cluster centers belong, form a cluster for each cluster center and all non-cluster center industry chain nodes associated therewith;

[0024] S428. Continuously repeat the steps S423 to S427 until the change distance of all new cluster centers is not less than a preset threshold when determining the new cluster centers of all clusters, and then take all new cluster centers and the non-cluster center industry chain nodes associated therewith as the final clustering result to form the corresponding categories.

[0025] Preferably, in the step S31, the cosine similarity algorithm is used to calculate the similarity between each enterprise in the park, and the specific calculation formula is as follows:

[0026]

[0027] Wherein, cos(θ) represents the similarity between different enterprises; d represents the dimension number of the corresponding enterprise basic information data; x i , y i respectively represent the i-th dimension data of different enterprises.

[0028] Preferably, in the step S31, if the similarity between two enterprises exceeds a certain threshold, a connection edge is established between the two network nodes corresponding to the two enterprises respectively, so that all network nodes and the connection edges established between the network nodes form an enterprise similarity network.

[0029] Preferably, in the step S31, the number of connection edges of each network node in the enterprise similarity network is analyzed to analyze the competition level of each enterprise in the park.

[0030] Preferably, in the step S43, the key industrial chain node comprehensive evaluation index is calculated according to the key industrial chain node comprehensive evaluation index of all industrial chain nodes in each category, the mean value of the enterprise monthly average power consumption overcapacity value, the mean value of the enterprise power expansion installation demand score, and the mean value of the daily average power consumption overcapacity value of the new park-entering enterprise, wherein the smaller the mean value of the enterprise competition level, the larger the mean value of the enterprise monthly average power consumption overcapacity value, the larger the mean value of the enterprise power expansion installation demand score, and the smaller the mean value of the daily average power consumption overcapacity value of the new park-entering enterprise, the larger the key industrial chain node comprehensive evaluation index of the corresponding category, and the more the industrial chain nodes in the category are identified as key industrial chain nodes.

[0031] The application further provides a computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the steps in the method for identifying key industrial chain nodes in a park based on multi-source data.

[0032] The application further provides a system for identifying key industrial chain nodes in a park based on multi-source data, comprising a processor and a computer readable storage medium connected to each other.

[0033] The application has the following beneficial effects: the application calculates a plurality of key identification indexes of each enterprise in a park according to multi-source data after fusion alignment processing, and then performs clustering analysis and key industrial chain node identification on the industrial chain nodes corresponding to each enterprise in the park according to the key identification indexes of each enterprise in the park, so that the key industrial chain nodes of the enterprises to be introduced into the park can be identified, which provides a reference for enterprise screening and introduction in the park marketing process, improves the adaptation degree of the park marketing target and the industrial chain in the park, and realizes the long-term stable and orderly development of the park. BRIEF DESCRIPTION OF DRAWINGS

[0034] Figure 1 It is a method flow diagram for identifying key industrial chain nodes in a park based on multi-source data. DETAILED DESCRIPTION

[0035] The application will be further described in detail below in combination with specific embodiments.

[0036] The embodiment provides a system for identifying key industrial chain nodes in a park based on multi-source data, which comprises a computer readable storage medium and a processor connected to each other, the computer readable storage medium has a computer program stored thereon, and the computer program is executed by the processor to implement the method for identifying key industrial chain nodes in a park based on multi-source data as shown in the figure. Figure 1 The method comprises the following steps S1, S2, S3 and S4.

[0037] S1. Obtain multi-source data of each enterprise in the park from different data systems, and perform fusion alignment processing on the multi-source data of each enterprise in the park respectively, wherein the multi-source data includes park enterprise list, enterprise basic information data and power contract data.

[0038] In this embodiment, information collection technology is used to obtain multi-source data of each enterprise in the park from different data systems. The multi-source data includes park enterprise list from the park management system, enterprise basic information data from the enterprise information query platform, and power contract data from the power system. The park enterprise list refers to a list recording the names of all enterprises in the park. The enterprise basic information data includes enterprise name, unified social credit code, business scope, and annual average electricity consumption, enterprise size, registration amount, and the number of insured persons that can reflect the nature of the enterprise. The power contract data includes electricity user archives, electricity charge and electricity quantity, and industry expansion installation data. It should be noted that industry expansion installation refers to the power grid system accepting the application of electricity users, and specifying a safe, economic and reasonable power supply scheme according to the electricity demand of electricity users and combining the condition of the power supply network. The industry expansion installation data includes grid rated power, electricity duration applied by electricity users, and electricity contract capacity applied by electricity users.

[0039] It should be noted that in different data systems, the same data of the same enterprise may have differences in recording methods, resulting in the situation that the same data between different data systems cannot be matched. For example, the enterprise name of an enterprise on the enterprise information query platform is different from the enterprise name of the enterprise on the power system. Therefore, the enterprise basic information data and the power contract data of the enterprise cannot be matched with each other. Therefore, after obtaining the multi-source data of each enterprise in the park from different data systems, semantic recognition technology is used to perform fusion alignment processing on the multi-source data of each enterprise in the park respectively, so that the same data of the same enterprise from different data systems can be matched with each other.

[0040] S2. Analyze the industry chain to which each enterprise in the park belongs and the industry chain node in the industry chain according to the multi-source data.

[0041] After the obtained multi-source data is fused and aligned, since the multi-source data includes enterprise basic information data, and the enterprise basic information data includes business scope information, the system can analyze the industrial chain to which each enterprise in the park belongs and the industrial chain node at which each enterprise in the park is located according to the business scope information of each enterprise in the park by using semantic recognition technology. For example, the business scope information of enterprise A contains the word "mobile phone sales", so the semantic recognition technology can analyze that enterprise A belongs to the mobile phone industrial chain, and enterprise A is at the industrial chain node corresponding to "mobile phone sales" in the mobile phone industrial chain; the business scope information of enterprise B contains the words "automobile manufacturing and sales", so the semantic recognition technology can analyze that enterprise B belongs to the automobile industrial chain, and enterprise B is at the industrial chain nodes corresponding to "automobile manufacturing" and "automobile sales" in the automobile industrial chain.

[0042] S3. Calculate a plurality of key identification indexes of each enterprise in the park according to the multi-source data, wherein the key identification indexes include the competition level degree of each enterprise in the park, the monthly average power consumption overcapacity value, the power consumption business expansion demand score, and the daily average power consumption overcapacity value of the new park-entering enterprise.

[0043] After the obtained multi-source data is fused and aligned, a plurality of key identification indexes of each enterprise in the park are calculated according to the multi-source data, wherein the key identification indexes are indexes used to identify the key industrial chain nodes to be introduced, and specifically include the competition level degree of each enterprise in the park, the monthly average power consumption overcapacity value, the power consumption business expansion demand score, and the daily average power consumption overcapacity value of the new park-entering enterprise. The calculation of the plurality of key identification indexes of each enterprise in the park specifically includes the following steps S31, S32, S33 and S34.

[0044] S31. Based on the complex network model, network nodes corresponding to each enterprise in the park are generated according to the park enterprise list, and then an enterprise similarity network is constructed according to the similarity between the network nodes and each enterprise in the park, and the competition level degree of each enterprise in the park is analyzed according to the enterprise similarity network, wherein the similarity between each enterprise in the park is specifically calculated according to the enterprise basic information data of each enterprise in the park.

[0045] The complex network theory proposes that there are various complex networks in nature and human society, and a complex network refers to a network that has some or all of the properties of self-organization, self-similarity, attractor, small world, and scale-free. A complex network model refers to a model that can construct a complex network based on the complex network theory. In this embodiment, based on the complex network model, an enterprise similarity network is constructed according to the park enterprise list and the similarity between each enterprise in the park, to analyze the competition level degree of each enterprise in the park, specifically including the following steps S311, S312 and S313.

[0046] S311. According to the park enterprise list, a plurality of network nodes respectively corresponding to each enterprise in the park are generated in the enterprise similarity network, so that the number of network nodes in the enterprise similarity network is equal to the number of enterprises in the park.

[0047] S312. The enterprise similarity network is constructed according to the similarity between the network nodes and each enterprise in the park.

[0048] In the embodiment, the similarity between each enterprise in the park is specifically calculated according to the enterprise basic information data of each enterprise in the park. Since the enterprise basic information data of each enterprise in the park includes enterprise name, unified social credit code, and annual average electricity consumption, enterprise size, registration amount, and the number of insured persons that can reflect the nature of the enterprise, the system can calculate the similarity between each enterprise in the park according to the annual average electricity consumption, enterprise size, registration amount, and the number of insured persons in the enterprise basic information data of each enterprise in the park, using the cosine similarity algorithm, and the specific calculation formula is as follows:

[0049]

[0050] Wherein, cos(θ) represents the similarity between different enterprises, and the greater the similarity means the greater the competition between different enterprises; d represents the dimension number of the corresponding enterprise basic information data, which is specifically 4 in the embodiment; x i , y i respectively represent the i-th dimension data of different enterprises, wherein the dimension data refers to annual average electricity consumption, enterprise size, registration amount, or the number of insured persons, for example, x i represents the i-th dimension data of enterprise A, and y i represents the i-th dimension data of enterprise B. Correspondingly, the first dimension data x1 of enterprise A refers to the annual average electricity consumption of enterprise A, the first dimension data y1 of enterprise B refers to the annual average electricity consumption of enterprise B, the second dimension data x2 of enterprise A refers to the enterprise size of enterprise A, the second dimension data y2 of enterprise B refers to the enterprise size of enterprise B, the third dimension data x3 of enterprise A refers to the registration amount of enterprise A, the third dimension data y3 of enterprise B refers to the registration amount of enterprise B, the fourth dimension data x4 of enterprise A refers to the number of insured persons of enterprise A, and the fourth dimension data y4 of enterprise B refers to the number of insured persons of enterprise B.

[0051] After the similarity between each enterprise in the park is calculated, if the similarity between two enterprises exceeds a certain threshold (for example, 0.7), it means that the two enterprises are similar enterprises, and thus a connection edge can be established between the two network nodes corresponding to the two enterprises. After the connection edges are established between the network nodes with a similarity exceeding a certain threshold according to the similarity between each enterprise in the park, all the network nodes and the connection edges established between the network nodes form the enterprise similarity network.

[0052] S313. Analyze the competition level of each enterprise in the park according to the enterprise similarity network.

[0053] In the enterprise similarity network, the number of connection edges of each network node reflects the number of competing enterprises of the enterprise corresponding to the network node in the park, that is, the number of connection edges of each network node reflects the competition level of the enterprise corresponding to the network node in the park. Therefore, the competition level of each enterprise in the park can be analyzed according to the number of connection edges of each network node in the enterprise similarity network. The greater the competition level of an enterprise, the higher the competition degree of the industry chain node where the enterprise is located in the park, and the less the enterprise needs to be introduced.

[0054] S32. Calculate the theoretical monthly average power consumption of each enterprise in the park according to the power contract data of each enterprise in the park, obtain the actual monthly average power consumption of each enterprise in the park, and compare the actual monthly average power consumption of each enterprise in the park with the theoretical monthly average power consumption to obtain the monthly average power consumption excess capacity value of each enterprise in the park.

[0055] Since the power contract data of each enterprise in the park includes electricity user files, electricity charges and power, industry expansion and installation data, and the industry expansion and installation data contains grid rated power, electricity application time of electricity users, and electricity contract capacity of electricity users, the theoretical monthly average power consumption of each enterprise in the park can be calculated according to the power contract data of each enterprise in the park, and the calculation formula is as follows:

[0056]

[0057] Wherein, W' represents the theoretical monthly average power consumption of the enterprise calculated according to the power contract data; P is the grid rated power in the power contract data; T is the electricity application time of the electricity user in the power contract data, the calculation unit is month; P max is the electricity contract capacity of the electricity user in the power contract data; is the power factor, which depends on the grid power factor standard. The power factor of industrial power is generally 0.85-0.9, and the specific value of this embodiment is 0.9.

[0058] Then the system obtains the actual power consumption of each enterprise in the park from the electricity charge and electricity quantity in the power contract data of each enterprise in the park, and calculates the actual monthly average power consumption W of each enterprise in the park in combination with the actual power consumption duration of each enterprise in the park, for example, the actual power consumption is W0, and the actual power consumption duration is T0, then the actual monthly average power consumption is W0 / T0. Then the system obtains the actual power consumption of each enterprise in the park from the electricity charge and electricity quantity in the power contract data of each enterprise in the park, and calculates the actual monthly average power consumption W of each enterprise in the park in combination with the actual power consumption duration of each enterprise in the park, for example, the actual power consumption is W0, and the actual power consumption duration is T0, then the actual monthly average power consumption is W0 / T0.

[0059] The monthly average power consumption capacity value η of each enterprise in the park reflects the power consumption usage state of each enterprise in the park, and the greater the monthly average power consumption capacity value η of the enterprise, the greater the market space of the industry chain node where the enterprise is located in the park, and the more the enterprise needs to be introduced.

[0060] S33. According to the power consumption and annual maximum load of each enterprise in the park within a certain time range, the predicted power consumption and predicted annual maximum load of each enterprise in the park within a future preset time are predicted, and the prediction results are compared with the current annual power consumption and current annual maximum load, and the power expansion demand score of each enterprise in the park is analyzed according to the comparison result.

[0061] The power expansion demand of the enterprise is an important measure of the future power demand of the enterprise, and the future power demand of each enterprise in the park can well reflect the production demand of the enterprise, so this embodiment analyzes the development prospect of the park by predicting the power expansion demand of each enterprise in the park, which specifically includes the following steps S331 and S332.

[0062] S331. Based on the gray Verhulst model, the predicted power consumption and predicted annual maximum load of each enterprise in the park within a future preset time (for example, 3 years in the future) are predicted.

[0063] Firstly, the power consumption and annual maximum load of each enterprise in the park within a certain time range (for example, in the past 20 years) are obtained, and then the original data sequence is established based on the power consumption and annual maximum load of each enterprise in the park within a certain time range. It should be noted that the original data sequence established based on the power consumption is the same as the original data sequence established based on the annual maximum load, which is as follows:

[0064] X 0 ={x 0 (1),x 0 (2),x 0 (3),...,x 0 (u)};

[0065] X 0represents the original data sequence established on the basis of the power consumption or annual maximum load of each enterprise in the park within a certain time range or in the past years; x 0 (1) represents the power consumption or annual maximum load in the year before the first year, x 0 (2) represents the power consumption or annual maximum load in the year before the second year, x 0 (3) represents the power consumption or annual maximum load in the year before the third year, …, x 0 (u) represents the power consumption or annual maximum load in the year before the u-th year, and in the present embodiment, u is specifically 20.

[0066] In the second step, the data in the original data sequence is accumulated and a close-to-mean sequence is generated.

[0067] The formula for accumulating the data in the original data sequence is as follows:

[0068]

[0069] wherein, x 0 (t') represents the power consumption or annual maximum load of each enterprise in the park in the year before t'; x (1) (t) represents the accumulation of the power consumption or annual maximum load of each enterprise in the park in the year before t, t = 1, 2, 3, …, u, for example:

[0070] t = 1,

[0071] t = 2,

[0072] t = 3,

[0073]

[0074] t = u,

[0075] Then, the accumulated sequence X (1) is generated according to the above-accumulated results, specifically as follows:

[0076]

[0077] Then, the close-to-mean sequence Z (1) of the accumulated sequence X (1) is generated, specifically as follows:

[0078] Z (1) = {z (1) (2), z (1) (3), z (1) (4), …, z (1) (t)};

[0079] in:

[0080] This represents the average of the electricity consumption or maximum annual load in the second year, calculated cumulatively from the data of two adjacent years.

[0081] This represents the average of the electricity consumption or maximum annual load in the third year, calculated cumulatively from the data of two adjacent years.

[0082] This represents the average of the electricity consumption or maximum annual load in the fourth year, calculated cumulatively from the data of two adjacent years.

[0083] ...

[0084] This represents the average of the electricity consumption or maximum annual load in year t, calculated by accumulating the data from two adjacent years.

[0085] The third step is to establish a grey Verhulst model, and based on the grey Verhulst model, predict the expected electricity consumption or the expected maximum annual load of each enterprise in the park within a preset time period.

[0086] Using the data from the original data sequence and the nearest mean sequence Z (1) The function that generates the grey Verhulst model from the data is as follows:

[0087] x (0) (t)+αz (1) (t)=β(z (1) (t)) 2 ;

[0088] Where α and β are undetermined parameters.

[0089] The differential equation of the gray Verhulst model function is:

[0090] Where t is time.

[0091] The solution to the differential equation is:

[0092]

[0093] Based on the solution of the differential equation, the time response sequence of the grey Verhulst model is obtained as follows:

[0094]

[0095] Because x (1) It is made by x (0) It is obtained by accumulation, so it can be obtained by adjusting x.(1) (t+1) is subtracted cumulatively to obtain x (0) The predicted value, i.e., the grey Verhulst model, is:

[0096] x (0) (t+1)=x (1) (t+1)-x (1) (t).

[0097] Then, the grey Verhulst model can be used to predict the expected electricity consumption and expected maximum annual load of each enterprise in the park within a predetermined time period (e.g., the next 3 years).

[0098] It should be noted that the undetermined parameters α and β are specifically calculated using the least squares method. Specifically, by substituting the data of t = 2, 3, ..., u into the above gray Verhulst model function, we obtain:

[0099]

[0100] The gray Verhulst model is written in matrix form, specifically:

[0101] Let the data vector Y = (x (0) (2),x (0) (3),...,x (0) (u)) T ;

[0102] Let the data matrix

[0103] Let the parameter vector μ = (αβ) T ;

[0104] The grey Verhulst model can then be represented as Y = Bμ. The undetermined parameters α and β can be solved using the least squares method:

[0105] [α,β]=(B T B) -1 B T Y.

[0106] S332. Compare the prediction results of the grey Verhulst model with the current annual electricity consumption and the current annual maximum load, and analyze the electricity expansion application demand score of each enterprise in the park based on the comparison results.

[0107] In this embodiment, the main indicators used to evaluate the degree of business expansion application demand of enterprises are the average annual electricity consumption growth rate and the average annual maximum load growth rate. The average annual electricity consumption growth rate is obtained by comparing the projected electricity consumption predicted by the grey Verhulst model with the current annual electricity consumption. The average annual maximum load growth rate is obtained by comparing the projected annual maximum load predicted by the grey Verhulst model with the current annual maximum load. The system can obtain the current annual electricity consumption and current annual maximum load of each enterprise in the park based on the electricity cost and electricity consumption data from their respective power contract data. In addition, two other indicators are used to evaluate the degree of business expansion application demand of enterprises: the current load rate and the projected load rate for the next three years. The current load rate is obtained by comparing the enterprise's annual maximum load with its power contract capacity, and the projected load rate for the next three years is obtained by comparing the enterprise's projected annual maximum load for the next 1, 2, and 3 years with its power contract capacity. The specific calculation formulas for these four indicators are as follows:

[0108]

[0109] Among them, G rate1 It is the company's average annual electricity consumption growth rate, G rate2 It is the company's average annual maximum load growth rate; L rate1 It is the current load factor of the enterprise, L rate2 Q0 represents the company's projected load factor for the next three years; Q1, Q2, and Q3 are the company's current annual electricity consumption, and Q1, Q2, and Q3 are the company's projected electricity consumption for the next 1, 2, and 3 years, respectively, predicted by the grey Verhulst model; P0 is the company's current maximum annual load, and P1, P2, and P3 are the company's projected maximum annual load for the next 1, 2, and 3 years, respectively, predicted by the grey Verhulst model; P max This is the company's current electricity contract capacity.

[0110] The annual average electricity consumption growth rate G of the enterprise was calculated. rate1 Average annual maximum load growth rate G rate2 Current load rate L rate1 The projected load factor L over the next 3 years rate2 After these four indicators, the entropy weight method is used to perform a weighted comprehensive calculation to obtain the enterprise's electricity expansion application demand score. The specific steps are as follows:

[0111] The first step is to analyze the average annual electricity consumption growth rate G. rate1 Average annual maximum load growth rate G rate2 Current load rate L rate1 The projected load factor L over the next 3 years rate2These four indicators are standardized by first generating an indicator value sequence X = {x1, x2, x3, x4} based on their specific values; where x1, x2, x3, and x4 represent the average annual electricity consumption growth rate G. rate1 Average annual maximum load growth rate G rate2 Current load rate L rate1 The projected load factor L over the next 3 years rate2 The index value.

[0112] Then, standardization is performed, using the following formula:

[0113]

[0114] Among them, Z j This represents the result of standardizing the values ​​of each indicator; X j Let X represent the j-th index value in the index value sequence X, where j = 1, 2, 3, 4; min{X} represents the minimum index value in the index value sequence X, and max{X} represents the maximum index value in the index value sequence X.

[0115] The second step is to calculate the weight P of each indicator value. j The specific formula is as follows

[0116] The third step is to calculate the standard information entropy E of each indicator value. j The specific formula is as follows

[0117] The fourth step is to calculate the information utility value D of each indicator. j The specific formula is D. j =1-E j .

[0118] The fifth step is to normalize the information utility values ​​of each indicator to obtain the entropy weight W of each indicator. j The specific formula is as follows

[0119] Substituting j = 1, 2, 3, 4 into the formula in step 5, we can calculate the annual average electricity consumption growth rate G of the first indicator. rate1 The entropy weight W1, the second indicator is the average annual maximum load growth rate G rate2 The entropy weight W2, the third indicator is the current load rate L rate1 The entropy weight W3, the fourth indicator is the projected load factor L over the next 3 years. rate2 The entropy weight W4 is used to calculate the enterprise's electricity expansion application demand score F = W1 × G. rate1 +W2×G rate2 +W3×L rate1 +W4×L rate2.

[0120] The score of a company’s electricity expansion application needs reflects the future development prospects of each company in the park. The higher the score of a company’s electricity expansion application needs, the better the future development prospects of the industrial chain node in which the company is located in the park, and the more necessary it is to introduce companies.

[0121] S34. Identify new entrants to the park based on the list of enterprises in the park and the power contract data of each enterprise in the park, calculate the theoretical average daily electricity consumption of the new entrants based on the power contract data, obtain the actual average daily electricity consumption of the new entrants, and compare the actual average daily electricity consumption of the new entrants with the theoretical average daily electricity consumption to obtain the overcapacity value of the average daily electricity consumption of the new entrants. Specifically, this includes the following steps S341, S342 and S343.

[0122] S341. First, connect the power contract data of each enterprise in the park according to the list of enterprises in the park, obtain the starting time of power supply from the power grid system to each enterprise in the park, regard the starting time of power supply as the enterprise's entry time in the park, and identify enterprises whose entry time is within the most recent 3 months as new enterprises in the park.

[0123] S342. Calculate the theoretical average daily electricity consumption of newly established enterprises based on the electricity contract data of each enterprise in the park. The specific calculation formula is as follows:

[0124]

[0125] Where w′ represents the theoretical average daily electricity consumption of newly established enterprises calculated based on electricity contract data; P is the rated power of the power grid in the electricity contract data; T′ is the electricity usage duration applied for by electricity users in the electricity contract data, calculated in days; P max The electricity contract capacity applied for by electricity users in the electricity contract data; The power factor depends on the power grid power factor standard. The power factor for industrial electricity is generally 0.85 to 0.9, and the specific value in this embodiment is 0.9.

[0126] S343. The system obtains the actual electricity consumption of newly entered enterprises based on the electricity bill and electricity consumption data in the electricity contract data, and calculates the actual average daily electricity consumption w of the newly entered enterprises in combination with the entry time of the enterprises. For example, if the actual electricity consumption is w0 and the actual electricity consumption duration is T0′, then the actual average monthly electricity consumption is... Then, the actual monthly average electricity consumption w of each enterprise in the park is compared with the theoretical monthly average electricity consumption w′ to obtain the daily average electricity consumption overcapacity value λ of the newly entered enterprises. The specific formula is as follows:

[0127] The daily average electricity consumption overcapacity value λ of newly established enterprises reflects the difficulty of entering the industrial chain node where the new enterprise is located. The smaller the daily average electricity consumption overcapacity value λ of newly established enterprises, the easier it is for the industrial chain node where the new enterprise is located to enter the park, and the more necessary it is to introduce enterprises.

[0128] S4. Based on the key identification indicators of each enterprise in the park, conduct cluster analysis and identify key industrial chain nodes of each enterprise in the park.

[0129] After calculating the competitiveness level, monthly average electricity consumption overcapacity, electricity expansion application score, and daily average electricity consumption overcapacity of newly entered enterprises in the park, the system performs cluster analysis and key industrial chain node identification on the industrial chain nodes of each enterprise in the park based on these key identification indicators, specifically including the following steps S41, S42 and S43.

[0130] S41. Calculate the average values ​​of key identification indicators for each node in the industrial chain within the park, specifically including the following steps S411, S412, S413 and S414.

[0131] S411. Calculate the average competitiveness level of all enterprises corresponding to each node in the industrial chain within the park. The specific calculation formula is as follows:

[0132]

[0133] in, C represents the mean degree of competition among all enterprises belonging to the m-th node of the industrial chain within the park, where m = 1, 2, 3, ..., k; k represents the number of nodes in the industrial chain within the park; mi Let i represent the degree of competition of the i-th enterprise in the m-th industrial chain node within the park, where i = 1, 2, 3, ..., n; and n represent the number of enterprises in the park belonging to the m-th industrial chain node.

[0134] S412. Calculate the average monthly electricity consumption overcapacity of all enterprises corresponding to each node of the industrial chain within the park. The specific calculation formula is as follows:

[0135]

[0136] in, η represents the average monthly excess electricity consumption of all enterprises belonging to the m-th industrial chain node within the park, where m = 1, 2, 3, ..., k; k represents the number of industrial chain nodes within the park; mi Let represent the monthly average electricity consumption excess value of the i-th enterprise in the m-th industrial chain node within the park, where i = 1, 2, 3, ..., n; and n represent the number of enterprises belonging to the m-th industrial chain node within the park.

[0137] S413. Calculate the average score of electricity expansion application demand for all enterprises corresponding to each industrial chain node in the park. The specific calculation formula is as follows:

[0138]

[0139] in, F represents the average score of electricity expansion application demand for all enterprises belonging to the m-th industrial chain node within the park, where m = 1, 2, 3, ..., k; k represents the number of industrial chain nodes within the park; mi Let i represent the electricity expansion application score of the i-th enterprise in the m-th industrial chain node within the park, where i = 1, 2, 3, ..., n; and n represent the number of enterprises in the m-th industrial chain node within the park.

[0140] S414. Calculate the average daily electricity consumption overcapacity of newly entered enterprises corresponding to each node of the industrial chain within the park. The specific calculation formula is as follows:

[0141]

[0142] in, λ represents the average daily excess capacity of electricity consumption of newly established enterprises belonging to the m-th industrial chain node within the park, where m = 1, 2, 3, ..., k; k represents the number of industrial chain nodes within the park; mi Let represent the daily average electricity consumption excess value of the i-th new enterprise in the m-th industrial chain node within the park, where i = 1, 2, 3, ..., h; and h represent the number of new enterprises in the m-th industrial chain node within the park.

[0143] S42. Based on the average values ​​of key identification indicators of each enterprise in the park, cluster analysis is performed on the industrial chain nodes to which each enterprise belongs in the park, thereby dividing all industrial chain nodes into a preset number of different categories. Specifically, this includes the following steps: S421, S422, S423, S424, S425, S426, S427 and S428.

[0144] S421. Based on business needs, first determine the number of clusters required for all industry chain nodes. In this embodiment, the number is 4. Each category corresponds to a cluster formed by industry chain nodes.

[0145] S422. Randomly define the cluster centers of the four categories at four different industry chain nodes, and name each cluster center O. m Based on the key identification indicators of each cluster center, the data coordinates of each cluster center are set as follows: Where m takes the value of 1, 2, 3, ..., k, which is a random selection of four terms, depending on the four different industry chain nodes randomly defined for the cluster centers of the four categories. k represents the number of industry chain nodes within the park. For example, if the cluster centers of the four categories are randomly defined as the 1st, 5th, 8th, and 10th industry chain nodes, then the values ​​of m are 1, 5, 8, and 10, and the data coordinates of the four cluster centers are respectively...

[0146] Each non-cluster core industry chain node is named Q. r Based on the key identification indicators of each non-cluster core industry chain node, the data coordinates of each non-cluster core industry chain node are set as follows: Where r takes the value 1, 2, 3, ..., k and any number of terms after excluding the number of terms of the four cluster centers, and k represents the number of nodes in the industrial chain within the park.

[0147] S423. Define the continuous classification factor distance from non-cluster-center industry chain nodes to each cluster center as d(Q r O m Specifically, the Euclidean distance algorithm is used to calculate the specific value, and the calculation formula is as follows:

[0148] As mentioned above, after randomly defining the 1st, 5th, 8th, and 10th industry chain nodes as four cluster centers, the continuous classification factor distance from the r-th industry chain node (a non-cluster center) to the 1st industry chain node (a cluster center) is: The distance to the continuous classification factor of the 5th industrial chain node, which is the cluster center, is The distance to the continuous classification factor of the 8th industrial chain node, which is the cluster center, is The distance to the continuous classification factor of the 10th industrial chain node, which is the cluster center, is

[0149] For example:

[0150] As the second non-cluster-center node in the industry chain, its continuous classification factor distance to the first cluster-center node in the industry chain is: The distance to the continuous classification factor of the 5th industrial chain node, which is the cluster center, is The distance to the continuous classification factor of the 8th industrial chain node, which is the cluster center, is The distance to the continuous classification factor of the 10th industrial chain node, which is the cluster center, is

[0151] As the third non-cluster-center node in the industry chain, its continuous classification factor distance to the first cluster-center node in the industry chain is: The distance to the continuous classification factor of the 5th industrial chain node, which is the cluster center, is The distance to the continuous classification factor of the 8th industrial chain node, which is the cluster center, is The distance to the continuous classification factor of the 10th industrial chain node, which is the cluster center, is

[0152] ...

[0153] The k-th non-cluster-center node in the supply chain has a continuous classification factor distance to the 1-th cluster-center node. The distance to the continuous classification factor of the 5th industrial chain node, which is the cluster center, is The distance to the continuous classification factor of the 8th industrial chain node, which is the cluster center, is The distance to the continuous classification factor of the 10th industrial chain node, which is the cluster center, is

[0154] S424. Define the discrete classification factor distance from non-cluster-center industry chain nodes to each cluster center as b(Q). r O m Specifically, the difference degree is used to represent the discrete classification factor distance. The smaller the difference degree, the smaller the discrete classification factor distance. The calculation principle is to count the number of key identification indicators that are different between non-cluster core industry chain nodes and cluster cores. This number is the difference degree, which represents the discrete classification factor distance from non-cluster core industry chain nodes to cluster cores.

[0155] S425. Calculate the Q of non-clustered industry chain nodes based on continuous and discrete classification factor distances. r To each cluster center O m Distance Z r,m The specific calculation formula is Z. r,m =d(Q r O m )+b(Q r O m ).

[0156] For example:

[0157] As the second node in the supply chain that is not the cluster center, its distance to the first node in the supply chain that is the cluster center is Z. r,m =Z 2,1 =d(Q2,O1)+b(Q2,O1), the distance to the continuous classification factor of the 5th industrial chain node that serves as the cluster center is Z. r,m =Z 2,5 =d(Q2,O5)+b(Q2,O5), the distance to the continuous classification factor of the 8th industrial chain node that serves as the cluster center is Z. r,m =Z2,8 =d(Q2,O8)+b(Q2,O8), the distance to the continuous classification factor of the 10th industrial chain node (which is the cluster center) is Z. r,m =Z 2,10 =d(Q2,O 10 )+b(Q2,O 10 ).

[0158] As the third node in the supply chain that is not the cluster center, its distance to the first node in the supply chain that is the cluster center is Z. r,m =Z 3,1 =d(Q3,O1)+b(Q3,O1), where Z is the distance to the continuous classification factor of the 5th industry chain node that serves as the cluster center. r,m =Z 3,5 =d(Q3,O5)+b(Q3,O5), the distance to the continuous classification factor of the 8th industrial chain node that is the cluster center is Z. r,m =Z 3,8 =d(Q3,O8)+b(Q3,O8), the distance to the continuous classification factor of the 10th industrial chain node (which is the cluster center) is Z. r,m =Z 3,10 =d(Q3,O 10 )+b(Q3,O 10 ).

[0159] ...

[0160] The distance from the k-th non-cluster center node to the 1st cluster center node in the supply chain is Z. r,m =Z k,1 =d(Q k ,O1)+b(Q k The distance from O1) to the continuous classification factor of the 5th industrial chain node, which is the cluster center, is Z. r,m =Z k,5 =d(Q k ,O5)+b(Q k The distance from O5) to the continuous classification factor of the 8th industrial chain node, which is the cluster center, is Z. r,m =Z k,8 =d(Q k ,O8)+b(Q k The distance from O8) to the continuous classification factor of the 10th industrial chain node, which is the cluster center, is Z. r,m =Z k,10 =d(Q k O 10 )+b(Q k O 10 ).

[0161] S426. Find each non-cluster core chain node Q.r To each cluster center O m Distance Z r,m The minimum value among them, then Q of each non-cluster core industry chain node r Associated with the nearest cluster center O m In terms of the category to which it belongs.

[0162] For example:

[0163] For the second industry chain node that is not the cluster center, calculate its distance Z to the first industry chain node that is the cluster center. 2,1 The distance Z of the continuous classification factor to the 5th industrial chain node that serves as the cluster center. 2,5 The distance to the continuous classification factor of the 8th industrial chain node, which is the cluster center, is Z. 2,8 The distance Z of the continuous classification factor to the 10th industrial chain node, which is the cluster center. 2,10 Next, find the distance Z from the second non-cluster center node in the supply chain to each cluster center. 2,1 Z 2,5 Z 2,8 Z 2,10 The minimum value min{Z 2,1 Z 2,5 Z 2,8 Z 2,10 Assume the minimum value is Z. 2,8 That is, the cluster center closest to the second industry chain node is the eighth industry chain node, and then the second industry chain node, which is not a cluster center, is associated with the category to which the eighth industry chain node, the closest cluster center, belongs.

[0164] For the third industry chain node that is not the cluster center, calculate its distance Z to the first industry chain node that is the cluster center. 3,1 The distance Z of the continuous classification factor to the 5th industrial chain node that serves as the cluster center. 3,5 The distance to the continuous classification factor of the 8th industrial chain node, which is the cluster center, is Z. 3,8 The distance Z of the continuous classification factor to the 10th industrial chain node, which is the cluster center. 3,10 Next, find the distance Z from the second non-cluster center node in the supply chain to each cluster center. 3,1 Z 3,5 Z 3,8 Z 3,10 The minimum value min{Z 3,1 Z 3,5 Z 3,8 Z 3,10 Assume the minimum value is Z. 3,10That is, the cluster center closest to the 3rd industry chain node is the 10th industry chain node, and then the 3rd industry chain node, which is not a cluster center, is associated with the category to which the 10th industry chain node, which is closest to the cluster center, belongs.

[0165] ...

[0166] For the k-th supply chain node that is not the cluster center, calculate its distance Z to the 1-th supply chain node that is the cluster center. k,1 The distance Z of the continuous classification factor to the 5th industrial chain node that serves as the cluster center. k,5 The distance to the continuous classification factor of the 8th industrial chain node, which is the cluster center, is Z. k,8 The distance Z of the continuous classification factor to the 10th industrial chain node, which is the cluster center. k,10 Next, find the distance Z from the second non-cluster center node in the supply chain to each cluster center. k,1 Z k,5 Z k,8 Z k,10 The minimum value min{Z k,1 Z k,5 Z k,8 Z k,10 Assume the minimum value is Z. k,1 That is, the cluster center closest to the kth industry chain node is the 1st industry chain node, and then the kth industry chain node, which is not a cluster center, is associated with the category to which the 1st industry chain node, the closest cluster center, belongs.

[0167] S427. After associating all non-cluster-center industry chain nodes with the category of the nearest cluster-center, each cluster-center and all its associated non-cluster-center industry chain nodes form a cluster. Since there are 4 cluster-centers, 4 clusters will be formed. Then, a new cluster-center is found based on all industry chain nodes in each cluster. Specifically, the average data coordinates of all industry chain nodes in each cluster are first calculated. Where L = 1, 2, 3, 4 This represents the average of the degree of competition among all nodes in the supply chain within the L-th cluster. This represents the average monthly electricity consumption overcapacity of all enterprises in the L-th cluster. This represents the average of the scores for the electricity expansion application needs of enterprises at all nodes in the L-th cluster. This represents the average daily electricity consumption overcapacity of newly established enterprises in all industry chain nodes within the Lth cluster; then, the distance from the data coordinates of each industry chain node in the cluster to this average data coordinate is calculated using the following formula: Where, q gLLet L be the data coordinates of the supply chain nodes in the Lth cluster; then, take the supply chain node with the smallest change distance as the new cluster center in the cluster.

[0168] For example:

[0169] In the first cluster, the cluster center is the 1st industry chain node, and the non-cluster center industry chain nodes associated with it are the 4th, 7th, 9th, ..., kth industry chain nodes. First, calculate the average data coordinates of all industry chain nodes in the first cluster. in, This represents the average of the mean competitiveness degree values ​​of the firms corresponding to the 4th, 7th, 9th, ..., kth nodes in the first cluster. This represents the average monthly electricity consumption overcapacity value of enterprises corresponding to the 4th, 7th, 9th, ..., kth nodes in the first cluster. This represents the average of the average scores of the enterprise electricity expansion application demand corresponding to the 4th, 7th, 9th, ..., kth nodes in the first cluster. This represents the average daily electricity consumption overcapacity of newly established enterprises corresponding to the 4th, 7th, 9th, ..., kth nodes in the first cluster. Then, the Euclidean distance algorithm is used to calculate the average distance from the data coordinates of each industry chain node in the first cluster to this data coordinate. The formula for calculating the change distance is: Where, q g1 Let g1 be the data coordinates of the supply chain nodes in the first cluster, with values ​​of 1, 4, 7, 9, ..., k. Then, the supply chain node with the smallest change distance is taken as the new cluster center in the first cluster. Assuming that the supply chain node with the smallest change distance is the 9th supply chain node, then the 9th supply chain node is taken as the new cluster center in the first cluster.

[0170] In the second cluster, the cluster center is the 5th industry chain node, and the non-cluster center industry chain nodes associated with it are the 6th, 11th, 12th, and so on. First, calculate the average data coordinates of all industry chain nodes in the second cluster. in, This represents the average of the mean competitiveness level scores of the firms corresponding to the 6th, 11th, 12th... nodes in the industry chain within the second cluster. This represents the average monthly electricity consumption overcapacity value of the enterprises corresponding to the 6th, 11th, 12th... nodes in the industrial chain within the second cluster. This represents the average of the average scores of the enterprise electricity expansion application demand corresponding to the 6th, 11th, 12th... nodes in the industrial chain within the second cluster. This represents the average daily electricity consumption overcapacity of newly established enterprises corresponding to the 6th, 11th, 12th... nodes in the industrial chain within the second cluster. Then, the Euclidean distance algorithm is used to calculate the average distance from the data coordinates of each industrial chain node in the second cluster to that data coordinate. The formula for calculating the change distance is: Where, q g2 Let g2 be the data coordinates of the supply chain nodes in the second cluster, with values ​​of 5, 6, 11, 12, ... Then, the supply chain node with the smallest change distance is taken as the new cluster center in the second cluster. Assuming that the supply chain node with the smallest change distance is the 12th supply chain node, then the 12th supply chain node is taken as the new cluster center in the second cluster.

[0171] In the third cluster, the cluster center is the 8th industry chain node. The non-cluster center industry chain nodes associated with it are the 2nd, 13th, 14th, and so on. First, calculate the average data coordinates of all industry chain nodes in the third cluster. in, This represents the average of the mean competitiveness level scores of the firms corresponding to the 2nd, 8th, 13th, 14th... nodes in the value chain within the third cluster. This represents the average monthly electricity consumption overcapacity value of enterprises corresponding to the 2nd, 8th, 13th, 14th... nodes in the industrial chain within the third cluster. This represents the average of the average scores of the enterprise electricity expansion application demand corresponding to the 2nd, 8th, 13th, 14th... nodes in the industrial chain within the third cluster. This represents the average daily electricity consumption overcapacity value of newly established enterprises corresponding to the 2nd, 8th, 13th, 14th... nodes in the industrial chain within the third cluster. Then, the Euclidean distance algorithm is used to calculate the average distance from the data coordinates of each industrial chain node in the third cluster to that data coordinate. The formula for calculating the change distance is: Where, q g3 Let g3 be the data coordinates of the supply chain nodes in the third cluster, with values ​​of 2, 8, 13, 14, ... Then, the supply chain node with the smallest change distance is taken as the new cluster center in the third cluster. Assuming that the supply chain node with the smallest change distance is the second supply chain node, then the second supply chain node is taken as the new cluster center in the third cluster.

[0172] In the fourth cluster, the cluster center is the 10th industry chain node. The non-cluster center industry chain nodes associated with it are the 3rd, 15th, 16th, and so on. First, calculate the average data coordinates of all industry chain nodes in the fourth cluster. in, This represents the average of the mean competitiveness level scores of the firms corresponding to the 3rd, 10th, 15th, 16th... nodes in the industry chain within the fourth cluster. This represents the average monthly electricity consumption overcapacity value of enterprises corresponding to the 3rd, 10th, 15th, 16th... nodes in the industrial chain within the fourth cluster. This represents the average of the average scores of the enterprise electricity expansion application demand corresponding to the 3rd, 10th, 15th, 16th... nodes in the industrial chain within the fourth cluster. This represents the average daily electricity consumption overcapacity value of newly established enterprises corresponding to the 3rd, 10th, 15th, 16th... nodes in the fourth cluster. Then, the Euclidean distance algorithm is used to calculate the average distance from the data coordinates of each industry chain node in the fourth cluster to that data coordinate. The formula for calculating the change distance is: Where, q g4 Let g4 be the data coordinates of the supply chain nodes in the fourth cluster, with values ​​of 3, 10, 15, 16, etc. Then, the supply chain node with the smallest change distance is taken as the new cluster center in the fourth cluster. Assuming that the supply chain node with the smallest change distance is the 10th supply chain node, the 10th supply chain node is still taken as the new cluster center in the fourth cluster.

[0173] S428. Repeat steps S423 to S427 until the distance between the four new cluster centers is not less than a preset threshold (e.g., 0.2) when determining the new cluster centers of the four clusters. Then, use the four new cluster centers and their associated non-cluster center industry chain nodes as the final clustering result to form four categories.

[0174] S43. After cluster analysis, obtain the mean values ​​of key identification indicators for all industrial chain nodes in each category, identify key industrial chain nodes for each category based on the results, and generate corresponding visual labels.

[0175] After clustering the industrial chain nodes within the park to obtain four categories, the average values ​​of key identification indicators for all industrial chain nodes in each category are obtained. Based on the results, key industrial chain nodes are identified for each category. It should be noted that the lower the average competitiveness score, the higher the average monthly electricity consumption overcapacity, the higher the average score for electricity expansion application demand, and the lower the average daily electricity consumption overcapacity of newly entered enterprises in a certain category, the higher the degree of need for enterprises to be introduced into the industrial chain nodes in that category, and the more likely the industrial chain nodes in that category should be identified as key industrial chain nodes. Specifically, in step S427 above, the average data coordinates of all industry chain nodes in a cluster formed by the cluster center and all its associated non-cluster center industry chain nodes have been calculated. This average data coordinates corresponds to the average values ​​of various key identification indicators for all industry chain nodes in the cluster. Therefore, after determining the new cluster centers of the four clusters and using the four new cluster centers and their associated non-cluster center industry chain nodes as the final clustering results to form four categories, the average data coordinates of all industry chain nodes in each cluster can be obtained as the average values ​​of various key identification indicators for all industry chain nodes in each category, specifically the average monthly electricity consumption overcapacity of enterprises. Average monthly electricity consumption exceeding capacity of enterprises Average score of enterprise electricity expansion application demand The average daily electricity consumption of newly established enterprises in the park exceeds the capacity limit. For the first category, the average values ​​of each key identification indicator are as follows: For the second category, the average values ​​of each key identification indicator were obtained as follows: For the third category, the average values ​​of each key identification indicator were obtained as follows: For the fourth category, the average values ​​of each key identification indicator were obtained as follows:

[0176] Then, based on the principle that the lower the average competitiveness score of enterprises in the industrial chain node, the higher the average monthly electricity consumption overcapacity of enterprises, the higher the average score of enterprises' electricity expansion application demand, and the lower the average daily electricity consumption overcapacity of newly entered enterprises, the higher the degree of enterprise introduction required in the industrial chain node of that category, and the more important it is to identify the industrial chain node in that category as a key industrial chain node, a comprehensive evaluation of key industrial chain nodes is carried out based on the average of various key identification indicators of all industrial chain nodes in each category. The specific formula is as follows:

[0177]

[0178] Where L = 1, 2, 3, 4; F LThe index represents the comprehensive evaluation index of key industrial chain nodes. The larger the value, the higher the degree to which enterprises need to be introduced into the industrial chain nodes in the Lth category, and the more important it is to identify the industrial chain nodes in the Lth category as key industrial chain nodes.

[0179] After calculating the comprehensive evaluation index F of key industrial chain nodes in each category L Then, key industrial chain nodes that need to be introduced can be identified. Based on the identification results of key industrial chain nodes, corresponding visual labels are generated for people to view, thereby providing a reference for the selection and introduction of enterprises in the park's investment promotion process, improving the compatibility between the park's investment promotion goals and the industrial chain within the park, so as to achieve the park's stable and orderly long-term development.

[0180] The above description is merely an embodiment of the present invention and does not limit the scope of patent protection. Any non-substantial changes or substitutions made by those skilled in the art based on the present invention will still fall within the scope of patent protection.

Claims

1. A method for identifying key industrial chain nodes within a park based on multi-source data, characterized in that, Includes the following steps: S1. Obtain multi-source data from various enterprises within the park from different data systems, and perform fusion and alignment processing on the multi-source data of various enterprises within the park respectively. The multi-source data includes a list of enterprises in the park, basic enterprise information data, and power contract data. S2. Analyze the industrial chain to which each enterprise in the park belongs and its position in the industrial chain based on the multi-source data; S3. Calculate multiple key identification indicators for each enterprise in the park based on the multi-source data. The key identification indicators include the competitiveness level of each enterprise in the park, the monthly average electricity consumption overcapacity value, the electricity expansion application demand score, and the daily average electricity consumption overcapacity value of newly entered enterprises. Specifically, it includes the following steps S31, S32, S33 and S34. S31. Based on a complex network model, network nodes corresponding to each enterprise in the park are generated according to the list of enterprises in the park. Then, an enterprise similarity network is constructed according to the similarity between the network nodes and each enterprise in the park. The competitive level degree of each enterprise in the park is analyzed according to the enterprise similarity network. The similarity between each enterprise in the park is specifically calculated based on the basic information data of each enterprise in the park. S32. Calculate the theoretical average monthly electricity consumption of each enterprise in the park based on the electricity contract data of each enterprise in the park, obtain the actual average monthly electricity consumption of each enterprise in the park, and compare the actual average monthly electricity consumption of each enterprise in the park with the theoretical average monthly electricity consumption to obtain the overcapacity value of the average monthly electricity consumption of each enterprise in the park. S33. Based on the electricity consumption or annual maximum load of each enterprise in the park within a certain time range, predict the expected electricity consumption and expected annual maximum load of each enterprise in the park within a future preset time. Compare the prediction results with the current annual electricity consumption and the current annual maximum load respectively. Analyze the electricity expansion application demand score of each enterprise in the park based on the comparison results. S34. Identify new enterprises entering the park based on the list of enterprises in the park and the power contract data of each enterprise in the park, calculate the theoretical average daily electricity consumption of the new enterprises based on the power contract data, obtain the actual average daily electricity consumption of the new enterprises, and compare the actual average daily electricity consumption of the new enterprises with the theoretical average daily electricity consumption to obtain the overcapacity value of the average daily electricity consumption of the new enterprises. S4. Based on the key identification indicators of each enterprise in the park, cluster analysis and key industrial chain node identification are carried out on the industrial chain nodes in which each enterprise in the park is located, specifically including the following steps S41, S42 and S43; S41. Calculate the average values ​​of key identification indicators for each node in the industrial chain; S42. Based on the average values ​​of key identification indicators of each enterprise in the park, cluster analysis is performed on the industrial chain nodes to which each enterprise belongs in the park, thereby dividing all industrial chain nodes into a preset number of different categories. S43. After cluster analysis, obtain the mean values ​​of key identification indicators for all industrial chain nodes in each category, identify key industrial chain nodes for each category based on the results, and generate corresponding visual labels. Step S42 includes: S421. Based on business needs, first estimate the number of clusters required for all nodes in the industry chain; S422. Randomly define the cluster centers of each category at different nodes in the industry chain; S423. Calculate the continuous classification factor distance from non-cluster core industry chain nodes to each cluster core; S424. Calculate the discrete classification factor distance from non-cluster core industry chain nodes to each cluster core; S425. Calculate the distance from non-cluster core industry chain nodes to each cluster core based on the continuous classification factor distance and the discrete classification factor distance; S426. Find the minimum distance from each non-cluster core industry chain node to each cluster core, and then associate each non-cluster core industry chain node with the category of the nearest cluster core; S427. After associating all non-cluster core industry chain nodes with the category to which the nearest cluster core belongs, each cluster core and all its associated non-cluster core industry chain nodes form a cluster; S428. Repeat steps S423 to S427 until, when determining the new cluster centers of all clusters, the change distance of all new cluster centers is not less than a preset threshold. Then, use all new cluster centers and their associated non-cluster center industry chain nodes as the final clustering result to form the corresponding categories.

2. The method for identifying key industrial chain nodes within a park based on multi-source data according to claim 1, characterized in that, In step S31, the cosine similarity algorithm is used to calculate the similarity between various enterprises in the park. The specific calculation formula is as follows: ; in, Indicates the similarity between different companies; This indicates the number of dimensions in the corresponding enterprise's basic information data; , These represent the i-th dimension data for different companies.

3. The method for identifying key industrial chain nodes within a park based on multi-source data according to claim 1, characterized in that, In step S31, if the similarity between two enterprises exceeds a certain threshold, a connection edge is established between the two network nodes corresponding to these two enterprises, so that all network nodes and the connection edges established between network nodes constitute an enterprise similarity network.

4. The method for identifying key industrial chain nodes within a park based on multi-source data according to claim 1, characterized in that, In step S31, the degree of competition level of each enterprise in the park is analyzed based on the number of connection edges of each network node in the enterprise similar network.

5. The method for identifying key industrial chain nodes within a park based on multi-source data according to claim 1, characterized in that, In step S43, the comprehensive evaluation index of key industrial chain nodes is calculated based on the comprehensive evaluation index of key industrial chain nodes in each category, the average monthly electricity consumption overcapacity of enterprises, the average score of enterprises' electricity expansion application, and the average daily electricity consumption overcapacity of newly entered enterprises. The smaller the average enterprise competition level score, the larger the average monthly electricity consumption overcapacity of enterprises, the larger the average score of enterprises' electricity expansion application, and the smaller the average daily electricity consumption overcapacity of newly entered enterprises, the larger the comprehensive evaluation index of key industrial chain nodes in the corresponding category, and the more likely the industrial chain nodes in that category should be identified as key industrial chain nodes.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps in the method for identifying key industrial chain nodes within a park based on multi-source data as described in any one of claims 1 to 5.

7. A system for identifying key industrial chain nodes within a park based on multi-source data, characterized in that, It includes interconnected processors and a computer-readable storage medium as described in claim 6.

Citation Information

Patent Citations

  • Industrial park agglomeration degree quantitative evaluation method based on multiple indicators

    CN103971000A

  • Industrial analysis method for data statistics based on industrial nodes

    CN112070345A