Industrial development path prediction method and device based on knowledge graph

Through deep learning algorithms, the economic development indicators are embedded encoding and sparsely processed, and the threshold is set to dynamically screen major industrial nodes, which solves the shortcomings of fixed threshold screening in the existing technology and improves the accuracy of industrial development path prediction.

CN119940638AInactive Publication Date: 2025-05-06INST OF GEOGRAPHY HENAN ACAD OF SCI

Patent Information

Application Number
CN202510051198.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the prior art, the use of fixed thresholds to screen major industrial nodes in the region is unable to fully adapt to the economic development level and diversity of industrial structures in different regions, resulting in a one-sided understanding of regional industrial characteristics and misjudgment of economic development trends.

Method used

By obtaining the economic development indicator set of the areas to be analyzed, using deep learning-based data processing algorithms for semantic embedding encoding, hub feature analysis and sparse processing, dynamically set thresholds to screen major industrial nodes, and thus predict industrial development paths.

Benefits of technology

Fully consider the economic development level and industrial structure diversity of different regions, dynamically adapt to industrial characteristics and development trends, and improve the accuracy of forecasting of industrial development paths.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940638A_ABST
    Figure CN119940638A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of industrial development path prediction, and particularly discloses an industrial development path prediction method and device based on a knowledge graph, and the method comprises the steps: firstly obtaining an economic development index set of a to-be-analyzed region, and introducing a deep learning-based data processing algorithm to carry out semantic embedded coding, hub feature analysis and sparse processing on each economic development index so as to mine the overall features and industry distribution rules of the regional economy, and then carrying out dynamic threshold setting by using a feedforward neural network model, screening out main industries of the region, and finally, carrying out industrial development on the region. And on this basis, development path prediction from the main industry to the target development industry is carried out. Through the mode, the economic development levels and the diversity of the industrial structure of different regions can be fully considered, and the industrial characteristics and the development trend of each region are dynamically adapted, so that the accuracy of industrial development path prediction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of industrial development path prediction, and more specifically, to a method and device for predicting industrial development paths based on a knowledge graph. Background Art

[0002] Traditionally, the prediction of industrial development paths relies on statistical analysis and expert experience. Although these methods can reflect the trend of industrial development to a certain extent, they are often limited by the comprehensiveness of data and the depth of analysis models, and it is difficult to accurately capture the complex and changeable correlations and dynamic evolution characteristics between industries. In recent years, with the rapid development of information technology, especially the widespread application of big data and artificial intelligence technology, data-driven industrial development path prediction methods have gradually become a research hotspot. As a powerful semantic network tool, knowledge graphs are widely used in knowledge management and intelligent decision support systems in various fields because they can effectively integrate and express complex domain knowledge.

[0003] For example, the invention patent with publication number CN112990575A proposes a method for predicting industrial development paths based on knowledge graphs. This method systematically collects industry and regional information to construct industry knowledge graphs and regional knowledge graphs. It not only takes into account the distribution of industries in geographical space, but also locates the main industrial nodes in each region through a threshold screening method based on the proportion of output value, constructs a regional industrial integration knowledge graph, and uses the industrial development scale index and development quality index to determine the target development industrial nodes. The shortest path algorithm is used to predict the optimal industrial development path from the main industrial nodes to the target development industrial nodes, thereby significantly improving the scientificity and accuracy of the industrial development path prediction and providing strong support for regional industrial planning.

[0004] However, the existing technology uses fixed thresholds to screen the main industrial nodes of a region. This static threshold screening mechanism may not be able to fully adapt to the diversity of economic development levels and industrial structures in different regions. Specifically, the economic environment and industrial development stage of different regions are often very different. Some regions may have highly concentrated dominant industries, while other regions may present a diversified industrial structure. In addition, the focus of industrial development in the same region in different historical periods may also change significantly. Therefore, the use of fixed threshold standards to screen major industrial nodes may lead to a one-sided understanding of regional industrial characteristics, or even misjudgment of regional economic development trends.

[0005] Therefore, we look forward to an optimized method and device for predicting industrial development paths based on knowledge graphs. Summary of the invention

[0006] In order to solve the above technical problems, the present application is proposed. The embodiment of the present application provides a method and device for predicting the industrial development path based on a knowledge graph, which first obtains a set of economic development indicators for the region to be analyzed, and introduces a data processing algorithm based on deep learning to perform semantic embedding coding, hub feature analysis and sparse processing on each economic development indicator, so as to dig out the overall characteristics of the regional economy and the law of industrial distribution, and then use a feedforward neural network model to set dynamic thresholds, screen out the main industries in the region, and then predict the development path from the main industry to the target development industry on this basis. In this way, the diversity of economic development levels and industrial structures in different regions can be fully considered, and the industrial characteristics and development trends of each region can be dynamically adapted to improve the accuracy of industrial development path prediction.

[0007] According to one aspect of the present application, a method for predicting industrial development paths based on a knowledge graph is provided, which includes: collecting massive amounts of industrial and regional information data to construct an industrial and regional information database; constructing an industrial knowledge graph based on the complete consumption coefficient between industries, and constructing a regional knowledge graph based on the geographical proximity between regions; taking industries whose output value share in the region is greater than a predetermined threshold as the main industrial nodes of the region, and constructing a regional industrial integration knowledge graph based on the output value share of the main industries in the region; calculating the development scale index and development quality index of all industries to determine the target development industrial node of any region in the regional industrial integration knowledge graph, and taking it as the end point, and taking any main industrial node of any region as the starting point, selecting the path with the largest complete consumption coefficient as the prediction of the optimal industrial development path of any region.

[0008] In the above-mentioned method for predicting industrial development paths based on knowledge graph, the setting of the preset threshold value includes:

[0009] Get a set of regional economic development indicators;

[0010] Embedding and coding each regional economic development indicator in the regional economic development indicator set to obtain a set of regional economic development indicator embedded coding vectors;

[0011] Performing a sparse processing on the set of regional economic development indicator embedded coding vectors based on feature distribution hubs to obtain a sparse set of regional economic development indicator embedded coding vectors;

[0012] The preset threshold is set based on the set of sparsely embedded coding vectors of regional economic development indicators.

[0013] According to another aspect of the present application, a device for predicting an industrial development path based on a knowledge graph is provided, which can implement the method for predicting an industrial development path based on a knowledge graph as described above, and includes:

[0014] The economic development indicator acquisition module is used to obtain the regional economic development indicator set;

[0015] An economic development indicator embedding coding module, used for embedding coding each regional economic development indicator in the regional economic development indicator set to obtain a set of regional economic development indicator embedding coding vectors;

[0016] A feature sparse processing module, used for performing a sparse processing on the set of regional economic development indicator embedded coding vectors based on a feature distribution hub to obtain a sparse set of regional economic development indicator embedded coding vectors;

[0017] The preset threshold setting module is used to set the preset threshold based on the set of sparsely embedded coding vectors of regional economic development indicators.

[0018] Compared with the prior art, the method and device for predicting the industrial development path based on the knowledge graph provided by the present application first obtains the set of economic development indicators of the region to be analyzed, and introduces a data processing algorithm based on deep learning to perform semantic embedding coding, hub feature analysis and sparse processing on each economic development indicator, so as to dig out the overall characteristics of the regional economy and the law of industrial distribution, and then use the feedforward neural network model to set dynamic thresholds, screen out the main industries in the region, and then predict the development path between the main industry and the target development industry on this basis. In this way, the diversity of economic development levels and industrial structures in different regions can be fully considered, and the industrial characteristics and development trends of each region can be dynamically adapted to improve the accuracy of industrial development path prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other purposes, features and advantages of the present application will become more apparent. The accompanying drawings are used to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0020] Figure 1 This is a flowchart of an industrial development path prediction method based on a knowledge graph according to an embodiment of the present application.

[0021] Figure 2 This is a data flow diagram of the industrial development path prediction method based on the knowledge graph according to an embodiment of the present application.

[0022] Figure 3 This is a flowchart of sub-step S3 of the method for predicting industrial development paths based on knowledge graph according to an embodiment of the present application.

[0023] Figure 4 This is a flowchart of sub-step S31 of the method for predicting industrial development paths based on knowledge graph according to an embodiment of the present application.

[0024] Figure 5 This is a flowchart of sub-step S32 of the method for predicting industrial development paths based on knowledge graph according to an embodiment of the present application. DETAILED DESCRIPTION

[0025] As shown in this application and claims, unless the context clearly indicates an exception, the words "a", "an", "an" and / or "the" do not refer to the singular and may also include the plural. Generally speaking, the terms "include" and "comprise" only indicate the inclusion of the steps and elements that have been clearly identified, and these steps and elements do not constitute an exclusive list. The method or device may also include other steps or elements.

[0026] Although the present application makes various references to certain modules in the system according to the embodiments of the present application, any number of different modules can be used and run on the user terminal and / or server. The modules are only illustrative, and different aspects of the system and method can use different modules.

[0027] Flowcharts are used in the present application to illustrate the operations performed by the system according to the embodiments of the present application. It should be understood that the preceding or following operations are not necessarily performed accurately in order. On the contrary, various steps may be processed in reverse order or simultaneously as required. Meanwhile, other operations may also be added to these processes, or a certain step or several steps of operations may be removed from these processes.

[0028] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described here.

[0029] It is worth noting that in this application, all actions to obtain data are carried out in compliance with the relevant data protection laws and policies of the country where the data is located, and with the authorization given by the owner of the corresponding device.

[0030] As mentioned in the background technology above, patent CN112990575A proposes a method for predicting industrial development paths based on knowledge graphs, which includes: collecting massive amounts of industrial and regional information data to build an industrial and regional information database; building an industrial knowledge graph based on the complete consumption coefficient between industries, and building a regional knowledge graph based on the geographical proximity between regions; taking industries whose output value accounts for a greater than a predetermined threshold in the region as the main industrial nodes of the region, and building a regional industrial integration knowledge graph based on the output value share of the main industries in the region; calculating the development scale index and development quality index of all industries to determine the target development industrial node of any region in the regional industrial integration knowledge graph, and taking it as the end point, and taking any main industrial node of any region as the starting point, selecting the path with the largest complete consumption coefficient as the prediction of the optimal industrial development path of any region.

[0031] Specifically, first, collect massive amounts of industry and regional information data and establish a comprehensive and detailed industry and regional information database to collect data that can reflect the overall picture of economic activities in a specific region, including but not limited to macroeconomic indicators, industry reports, corporate statements, academic research results, news information, and dynamic information on social media platforms.

[0032] After completing the data collection, the next step is to construct an industry knowledge graph based on the total consumption coefficient between industries (i.e., the degree of dependence of an industry on the products and services of other industries during the production process). This process first requires determining the input-output relationship between industries and quantifying the input-output relationship by calculating the total consumption coefficient. The total consumption coefficient reflects the total demand of an industry for the products of another industry in its entire production chain, including direct and indirect demand. By matrixing the total consumption coefficients between all industries, a complete industry association network can be obtained, in which each node represents an industry and the weight of the edge represents the intensity of dependence between the two industries. In order to improve the accuracy of the model, the influence of time factors, that is, the changes in the relationship between industries in different time periods, should also be considered. Therefore, when constructing the industry knowledge graph, a rolling window method should be used to regularly update the total consumption coefficient matrix to reflect the latest industry interactions. In addition, considering that some industries may have seasonal characteristics or be greatly affected by policies, additional adjustment parameters can be introduced to make the model more flexible and adaptable.

[0033] At the same time, it is also necessary to build a regional knowledge graph based on the geographical proximity between regions. This process is not just to simply connect each region as an independent node, but to deeply analyze the various forms of connection between regions, such as logistics, personnel flow, and capital flow. Specifically, the closeness between regions can be measured through multiple dimensions such as transportation network, communication infrastructure, and trade exchanges, and the weight of the edge can be set accordingly. For example, if there are frequent cargo transportation routes between two cities, it means that the two cities have a strong correlation in the industrial chain; and city pairs with frequent personnel flow often mean close cooperation in the service industry. In addition, long-term and stable regional cooperation relationships can be mined in combination with historical data to further enrich the content of the regional knowledge graph. When constructing a regional knowledge graph, it is also necessary to pay attention to changes in time and space. On the one hand, with the improvement of economic development level and adjustment of industrial structure, the relative status between different regions may change; on the other hand, new transportation projects, information technology development and other factors will also reshape the connection mode between regions. Therefore, a dynamic update mechanism should be maintained during the construction process to reflect the latest changes in a timely manner.

[0034] After completing the construction of the industry and regional knowledge graph, it is necessary to identify the main industry nodes in each region. Here, the method of using the output value ratio greater than the predetermined threshold is used to screen the main industry nodes. The core idea of ​​this method is to select those industries that play a key role in local economic growth as the main nodes based on the contribution of each industry to the total regional output value. In specific operations, the average annual output value ratio of each industry can be calculated based on historical statistical data, and a reasonable threshold (such as 5% or 10%) can be set. Any industry that exceeds this threshold is regarded as a main industry node.

[0035] In the above-mentioned industrial development path prediction method based on knowledge graph, by systematically collecting industry and regional information, constructing industrial knowledge graph and regional knowledge graph, not only the distribution of industry in geographical space is considered, but also the main industrial nodes of each region are located through the threshold screening method based on the output value ratio, and the regional industrial integration knowledge graph is constructed. The industrial development scale index and development quality index are used to determine the target development industrial node, and the shortest path algorithm is used to predict the optimal industrial development path from the main industrial node to the target development industrial node, thereby significantly improving the scientificity and accuracy of the industrial development path prediction. However, the above-mentioned method uses a fixed threshold to screen the main industrial nodes of the region. This static threshold screening mechanism may not be able to fully adapt to the diversity of economic development levels and industrial structures in different regions. Specifically, the economic environment and industrial development stage of different regions are often very different. Some regions may have highly concentrated dominant industries, while others may present a diversified industrial structure. In addition, the focus of industrial development in the same region in different historical periods may also change significantly. Therefore, the use of fixed threshold standards to screen the main industrial nodes may lead to a one-sided understanding of regional industrial characteristics, or even misjudgment of the development trend of the regional economy.

[0036] In response to the above technical problems, this application proposes an optimized method for predicting industrial development paths based on knowledge graphs, which first obtains a set of economic development indicators for the region to be analyzed, and introduces a data processing algorithm based on deep learning to perform semantic embedding coding, hub feature analysis, and sparse processing on each economic development indicator, so as to dig out the overall characteristics of the regional economy and the law of industrial distribution, and then use a feedforward neural network model to set dynamic thresholds and screen out the main industries in the region, so as to predict the development path from the main industry to the target development industry on this basis. In this way, the diversity of economic development levels and industrial structures in different regions can be fully considered, and the industrial characteristics and development trends of each region can be dynamically adapted to improve the accuracy of industrial development path prediction.

[0037] Figure 1 This is a flowchart of an industrial development path prediction method based on a knowledge graph according to an embodiment of the present application. Figure 2 Schematic diagram of data flow of the method for predicting industrial development paths based on knowledge graph according to an embodiment of the present application. Figure 1 and Figure 2As shown, the industrial development path prediction method based on knowledge graph includes the steps of: S1, obtaining a set of regional economic development indicators; S2, embedding and coding each regional economic development indicator in the set of regional economic development indicators to obtain a set of regional economic development indicator embedded coding vectors; S3, performing sparse processing on the set of regional economic development indicator embedded coding vectors based on feature distribution hubs to obtain a set of sparse regional economic development indicator embedded coding vectors; S4, setting the preset threshold based on the sparse set of regional economic development indicator embedded coding vectors.

[0038] In the above-mentioned industrial development path prediction method based on knowledge graph, the step S1 obtains a set of regional economic development indicators. It should be understood that economic development indicators are key information carriers that reflect the operation status of regional economy and the development trend of industry, covering data in multiple dimensions such as regional gross domestic product (GDP), GDP growth rate, specific operation data of various industries, unemployment rate, number of employees, number of enterprises, import and export trade volume, fixed asset investment, scientific and technological innovation input and output, etc., which can characterize the scale, structure, vitality and innovation ability of regional economy from different angles. For example, GDP intuitively shows the overall scale and growth trend of regional economy, and is an important comprehensive indicator for measuring regional economic strength; the specific operation data of various industries can reveal the internal operation efficiency and profitability of the industry, reflect the contribution share and development dynamics of various industries in the total economy, and help analyze the evolution of industrial structure; unemployment rate, number of employees and number of enterprises can reflect the employment absorption capacity of the industry and the activity of market entities; import and export trade volume reflects the outward orientation of regional economy and the degree of participation in the international market; fixed asset investment shows the long-term support for economic growth; indicators related to scientific and technological innovation input and output, such as R&D expenditure, patent applications and authorizations, highlight the innovation-driven potential of the industry. By comprehensively collecting these economic development indicators, we can build a multi-dimensional and comprehensive portrait of the regional economy, which will help us accurately grasp the overall characteristics of the regional economy and the patterns of industrial distribution, and set reasonable thresholds to screen out the main industrial nodes in the region.

[0039] First, in the process of obtaining regional economic development indicators, it is necessary to pay attention to the diversity of data sources. Usually, official statistics, industry reports and specific industry analysis reports provided by research institutions are combined, data is obtained directly from enterprises through multiple channels such as questionnaires or field interviews, and social media platforms, professional forums and technical blogs are used to capture the latest market dynamics and technology trends. The annual or quarterly economic reports released by government statistical departments provide the most authoritative macro data, such as gross domestic product (GDP), industrial added value, consumer price index (CPI), unemployment rate, etc.; while industry associations and research institutions can supplement the detailed analysis of specific industries, such as market share, import and export volume, average profit margin, etc. of each industry. In addition, by conducting questionnaires or field interviews on enterprises, first-hand enterprise-level data can be obtained, which helps to understand the actual operating conditions, challenges and future plans of enterprises. For emerging industries or innovative enterprises, non-traditional data sources such as social media platforms, professional forums and technical blogs can also be used to capture the latest market dynamics and technology trends.

[0040] After collecting the raw data, it is necessary to clean and preprocess the data to ensure the quality and consistency of the data. This step includes identifying and deleting duplicate data entries to ensure the uniqueness of each record; using statistical methods (such as mean and median filling) or machine learning algorithms to predict missing values ​​and improve data integrity; checking and correcting possible entry errors or other abnormalities; converting data from different sources into a unified format for subsequent processing and analysis; deflation of time series data to eliminate the impact of factors such as inflation and make data from different periods comparable; and selecting appropriate years or months as the analysis interval according to the purpose of the research to reflect short-term fluctuations or long-term trends.

[0041] In addition, in order to capture the complex connections and dynamic evolution characteristics between industries, the regional economic development indicator set can also include cross-industry interaction indicators. This includes but is not limited to upstream and downstream relationships in the supply chain, collaboration patterns in the industrial chain, and competitive situations within industrial clusters. Through the quantitative analysis of these non-traditional economic indicators, it is possible to reveal the subtle but crucial industrial linkage effects, thereby providing richer background information for predicting industrial development paths. For example, the development of the automobile industry not only depends on upstream raw material suppliers such as steel and rubber, but is also closely related to downstream companies such as electronic equipment manufacturers and financial service providers. Therefore, when constructing a set of regional economic development indicators, the key performance indicators (KPIs) of these related industries should be covered as much as possible, such as order volume, inventory turnover, financing scale, etc.

[0042] In the above-mentioned industrial development path prediction method based on knowledge graph, the step S2 embeds and codes each regional economic development indicator in the regional economic development indicator set to obtain a set of regional economic development indicator embedding coding vectors. Here, the present application takes into account that in the processing of economic data, the original economic indicator data is usually a mixed type of text and numerical data with semantic information, and the traditional statistical analysis method is difficult to directly capture the deep-level associations and potential laws between different indicator data. Therefore, in order to comprehensively capture the inherent semantic characteristics and relevance of each regional economic development indicator, the present application further embeds and codes each regional economic development indicator in the regional economic development indicator set to map it to a low-dimensional, dense vector space, and convert it into a vector representation containing rich semantic information, that is, a regional economic development indicator embedding coding vector. It should be understood that in a low-dimensional vector space, semantically similar or related economic development indicators will be mapped to a close position, thereby reflecting the economic relevance between the two. Through this encoding method, not only the semantic information of the original data is retained and the dimension of the data is reduced, but also the subsequent analysis process can better capture the potential connection between economic indicators, thereby improving the scientificity and accuracy of threshold setting.

[0043] In the above-mentioned industrial development path prediction method based on knowledge graph, the step S3 performs sparse processing on the set of regional economic development indicator embedded coding vectors based on feature distribution hubs to obtain a sparse set of regional economic development indicator embedded coding vectors. It should be understood that, considering that in actual economic data, the economic development indicator data is large in scale and contains a large amount of redundant information and noise. Therefore, in order to effectively mine the important features that have an impact on the screening of major industries, reduce computational complexity and improve prediction efficiency, the present application proposes a feature sparse processing method based on feature distribution hubs, which analyzes the overall feature distribution of regional economic development indicator data to identify and retain key features located at feature distribution hubs, while removing redundant or noise features, thereby constructing a streamlined and efficient set of economic development indicator features based on the overall economic characteristics of the region. Among them, Figure 3 Flow chart of sub-step S3 of the method for predicting industrial development paths based on knowledge graph according to an embodiment of the present application. Figure 3As shown, the step S3 includes the steps of: S31, performing a feature distribution hub search on the set of regional economic development indicator embedded coding vectors to obtain the regional economic development indicator embedded distribution hub coding vector; S32, based on the feature offset of each regional economic development indicator embedded coding vector in the set of regional economic development indicator embedded coding vectors relative to the regional economic development indicator embedded distribution hub coding vector, performing feature sparse processing on the set of regional economic development indicator embedded coding vectors to obtain the sparse set of regional economic development indicator embedded coding vectors.

[0044] Specifically, the step S31 performs a feature distribution hub search on the set of regional economic development indicator embedded coding vectors to obtain the regional economic development indicator embedded distribution hub coding vector. That is, by performing a feature distribution hub search on the set of regional economic development indicator embedded coding vectors, the key feature points at the distribution core or with obvious distinction are located to reflect the core elements and dominant trends of regional economic development. This method helps to identify the most representative economic characteristics and provide a basis for understanding the internal driving force and direction of regional economic growth. Figure 4 Flow chart of sub-step S31 of the method for predicting industrial development paths based on knowledge graph according to an embodiment of the present application. Figure 4 As shown, the step S31 includes the steps of: S311, calculating the characteristic difference factor of each regional economic development indicator embedded coding vector in the set of regional economic development indicator embedded coding vectors relative to the set of regional economic development indicator embedded coding vectors to obtain a set of regional economic development indicator semantic difference factors; S312, taking the inverse of each regional economic development indicator semantic difference factor in the set of regional economic development indicator semantic difference factors and performing normalization processing based on the Softmax function to obtain a set of regional economic development indicator hub feature semantic correlation factors; S313, using the set of regional economic development indicator hub feature semantic correlation factors as the weight distribution, weighted aggregation is performed on the set of regional economic development indicator embedded coding vectors to obtain the regional economic development indicator embedded distribution hub coding vector.

[0045] More specifically, in a specific example of the present application, the step S311 includes: respectively calculating the mean of the Mahalanobis distances of each regional economic development indicator embedded coding vector in the set of regional economic development indicator embedded coding vectors relative to other regional economic development indicator embedded coding vectors to obtain a set of semantic difference factors of the regional economic development indicators, which is expressed by the formula:

[0046] X={x1,x2,...,x i ,...,x n}

[0047]

[0048] Among them, X represents the set of embedding coding vectors of the regional economic development indicators, x1, x2, x i 、x j and x n denote the first, second, i-th, j-th and n-th regional economic development indicator embedded coding vectors in the set of regional economic development indicator embedded coding vectors, respectively, n is the number of regional economic development indicator embedded coding vectors, (·) denotes the transpose of the vector, S is the x i and the x j The covariance matrix between i For the x i Semantic differentiation factors of regional economic development indicators.

[0049] That is, by calculating the average distance between the embedded coding vectors of the economic development indicators of each region and the overall characteristics of the set, the correlation and hierarchical structure among the economic development indicators can be revealed, and the importance of the characteristic points of each economic development indicator can be preliminarily evaluated.

[0050] More specifically, the step S312 and the step S313 are expressed by the formula:

[0051]

[0052] Among them, exp(·) represents the exponential operation with a natural constant as the base, v h Represents the hub encoding vector of the embedded distribution of regional economic development indicators.

[0053] That is, firstly, through normalization processing, the weight distribution of each economic development indicator feature is generated to ensure the comparability between different indicators and highlight the relative importance of each feature in regional economic development. Then, through the weighted aggregation method, the feature distribution hub of the set of embedding coding vectors of the regional economic development indicators is constructed, that is, the embedding distribution hub coding vector of the regional economic development indicators is formed, which is used as the refinement and generalization expression of the core elements of regional economic development, and provides an important reference standard for the subsequent feature sparse processing. In this way, the key characteristics of the regional economy are effectively captured, laying the foundation for further analysis.

[0054] Figure 5 Flow chart of sub-step S32 of the method for predicting industrial development paths based on knowledge graph according to an embodiment of the present application. Figure 5As shown, the step S32 includes the steps of: S321, respectively calculating the semantic fine-grained ablation factors of each regional economic development indicator embedded coding vector in the set of regional economic development indicator embedded coding vectors relative to the regional economic development indicator embedded distribution hub coding vector to obtain a set of regional economic development indicator semantic fine-grained ablation factors; S322, based on the set of regional economic development indicator semantic fine-grained ablation factors, performing feature sparse processing on the set of regional economic development indicator embedded coding vectors to obtain the sparse set of regional economic development indicator embedded coding vectors.

[0055] More specifically, the step S321 is expressed by the formula:

[0056]

[0057] Among them, ||·|| represents the norm of the vector, max(·) represents the maximum value operation, ε is a very small positive number used to prevent the denominator from being zero, and a i Indicates that x i Semantic fine-grained ablation factors of regional economic development indicators.

[0058] That is, taking the embedding distribution hub coding vector of the regional economic development indicators as a benchmark, the contrastive learning method is used to calculate the semantic fine-grained ablation factor of the embedding distribution hub coding vector of each regional economic development indicator relative to the embedding distribution hub coding vector of the regional economic development indicators, so as to measure the degree of influence on the overall expression of the hub feature when removing the characteristics of the regional economic development indicators, thereby revealing the semantic association strength between each economic development indicator and the core elements of regional economic development. In this way, the importance and contribution of each indicator in describing the core characteristics of the regional economy can be more clearly understood.

[0059] More specifically, in a specific example of the present application, the step S322 includes: first, inputting each of the regional economic development indicator semantic fine-grained ablation factors in the set of regional economic development indicator semantic fine-grained ablation factors into a sparse module based on a gating function to obtain a set of sparse regional economic development indicator semantic fine-grained ablation factors, which is expressed by the formula:

[0060]

[0061] Where mask(·) represents the gating function, θ represents the gating threshold, and w i Indicates that x i The semantic fine-grained ablation factor of sparse regional economic development indicators.

[0062] Then, the set of the sparse regional economic development indicator semantic fine-grained ablation factors is used as the weight distribution, and the set of regional economic development indicator embedded coding vectors is weighted modulated to obtain the set of sparse regional economic development indicator embedded coding vectors, which is expressed by the formula:

[0063] X'={x i '}={w i *x i}

[0064] Among them, x i ' indicates that x i The corresponding sparse regional economic development indicator embedding coding vector, X' represents the set of sparse regional economic development indicator embedding coding vectors.

[0065] That is, the gating mechanism is further used to screen each semantic fine-grained ablation factor, and the regional economic development indicator embedding coding vector corresponding to the lower semantic fine-grained ablation factor is regarded as redundant or noise features, and filtered and eliminated. At the same time, the economic indicator features that have an important impact on the overall expression of the core elements of regional economic development are retained. This can effectively sparse the feature set of regional economic development indicators, so that the subsequent major industry screening process can focus more on key economic indicators, thereby improving analysis efficiency and reducing the waste of computing resources. In this way, the accuracy of feature selection can be ensured, and the explanatory power and practicality of the model can be enhanced.

[0066] In the above-mentioned industrial development path prediction method based on knowledge graph, the step S4 sets the preset threshold value based on the set of the sparse regional economic development indicator embedding coding vectors. In a specific example of the present application, the step S4 includes: inputting the set of the sparse regional economic development indicator embedding coding vectors into a threshold setter based on a feedforward neural network model to obtain the preset threshold value. Here, the powerful nonlinear fitting ability of the feedforward neural network is further utilized to determine the screening criteria for the main industries. In the technical solution of the present application, the feedforward neural network is composed of an input layer, a hidden layer and an output layer. By learning a large amount of historical data and known industrial classification cases, a complex mapping relationship between the characteristics of regional economic indicators and the threshold value of the output value of the main industries can be established. Among them, the input layer receives the set of regional economic development indicator embedding coding vectors after sparse processing, and the hidden layer extracts the deep-level economic indicator intrinsic correlation feature information through the connection and activation function between the multi-layer neurons, and constructs an internal representation that is instructive for the screening of the main industries. Finally, the output layer outputs a preset threshold value based on the output result of the hidden layer, which is used to determine which industries can be regarded as the main industrial nodes of the region. In this way, the output value share threshold can be dynamically adjusted according to the economic characteristics and data distribution of different regions, ensuring that the main industries screened out have high accuracy and representativeness, thereby providing strong support for the sustainable development of the regional economy and industrial upgrading.

[0067] In a preferred example of the present application, embedding the set of sparse regional economic development indicators into the coding vector through a threshold setter based on a feedforward neural network model to obtain a preset threshold includes:

[0068] Firstly, the set of the sparse regional economic development index embedding coding vectors is feature spliced ​​to obtain a regional economic development index embedding feature global splicing coding vector;

[0069] Secondly, the eigenvalues ​​of the regional economic development index embedding feature global concatenation coding vector are arranged in ascending order to form a regional economic development index embedding feature global concatenation sequence coding vector;

[0070] Next, in response to the absolute value of the difference between the ith eigenvalue and the i+1th eigenvalue of the global concatenated sequential coding vector of the regional economic development index embedded feature being less than or equal to the distance difference hyperparameter ε, the weighted sum between the ith eigenvalue and the i+1th eigenvalue is calculated as the optimized i+1th eigenvalue, and the square root of the sum of squares of all eigenvalues ​​of the global concatenated coding vector of the regional economic development index embedded feature is calculated, which is expressed by the formula:

[0071]

[0072] Among them, v i represents the ith eigenvalue of the global concatenated coding vector of the regional economic development indicator embedded features, L represents the length of the global concatenated coding vector of the regional economic development indicator embedded features, and γ1 represents the square root of the sum of the squares of all eigenvalues ​​of the global concatenated coding vector of the regional economic development indicator embedded features;

[0073] Then, multiply it by 2 and divide it by the square of the length of the regional economic development index embedding feature global splicing encoding vector to obtain the regional economic development index embedding feature global splicing space primitive value, which is expressed by the formula:

[0074] γ2=2×γ1 / L 2

[0075] Among them, γ2 represents the value of the global splicing space primitive of the regional economic development index embedding feature;

[0076] Next, in response to the absolute value of the difference between the i-th value and the i+1-th eigenvalue of the regional economic development indicator embedded feature global splicing sequential encoding vector being greater than the distance difference hyperparameter ε, after multiplying the regional economic development indicator embedded feature global splicing space primitive value by the i-th eigenvalue, a weighted subtraction between the product and the i+1-th eigenvalue is calculated to obtain the optimized i+1-th eigenvalue;

[0077] Then, on the basis of keeping the first eigenvalue of the regional economic development indicator embedded feature global splicing sequence coding vector unchanged, combining the optimized i+1th eigenvalue to obtain the optimized regional economic development indicator embedded feature global splicing coding vector;

[0078] Finally, the optimized regional economic development index is embedded in the feature global concatenation coding vector and passed through a threshold setter based on a feedforward neural network model to obtain a preset threshold.

[0079] Here, in the case where the set of regional economic development indicator embedded coding vectors represents the low-dimensional embedded coding features of regional economic development indicators, after the features are sparsed based on the feature distribution hubs, although the set of regional economic development indicator embedded coding vectors is subjected to feature selection and deletion, some complementary information will be lost, resulting in insufficient long-distance splicing coding representation of the global splicing coding vector of the regional economic development indicator embedded features, thereby reducing the expression effect of the global splicing coding vector of the regional economic development indicator embedded features and affecting the accuracy of the preset threshold obtained by the threshold setter based on the feedforward neural network model.

[0080] Therefore, the high-dimensional feature space primitive representation based on self-inner product fusion of the regional economic development indicator embedded feature global splicing coding vector is used to capture the complex structure of the global network interaction of its eigenvalues, thereby reconstructing the dynamic search representation relationship between the eigenvalues ​​of the regional economic development indicator embedded feature global splicing coding vector by simulating the scale-based high-dimensional feature space potential primitives, so as to realize the coding reconstruction of the real sequence distribution behavior of the regional economic development indicator embedded feature global splicing coding vector under long distance, improve the coding expression effect of the regional economic development indicator embedded feature global splicing coding vector, and improve the accuracy of the preset threshold obtained by the threshold setter based on the feedforward neural network model.

[0081] On this basis, based on the proportion of output value of major industries in the region, a regional industrial integration knowledge map is further constructed. This process involves analyzing the correlation between major industrial nodes and other industries in their regions, exploring the synergy and development potential between industries. Specifically, based on the proportion of output value of major industries in the region, its driving effect on surrounding industries can be evaluated, and a more complex regional industrial network can be constructed on this basis. For example, a major manufacturing node may attract a series of supporting service industries (such as logistics, finance, etc.) to gather, thus forming a fully functional industrial cluster. In this way, not only can the characteristics of the existing industrial layout be revealed, but also a reference basis can be provided for future industrial upgrading and structural adjustment.

[0082] In order to determine the target development industry node of any region in the regional industrial integration knowledge map, it is also necessary to calculate the development scale index and development quality index of all industries. The development scale index mainly measures the absolute scale of an industry, usually expressed by indicators such as output value, number of employees, and fixed asset investment; while the development quality index focuses on evaluating the efficiency and sustainability of the industry, such as profit margin, technological innovation capability, and environmental protection level. By comprehensively considering these two factors, the development status of each industry can be comprehensively evaluated, and the target nodes with the most development potential can be selected from them. Specifically, in terms of the calculation method, a weighted scoring method can be used to assign corresponding weights according to the importance of different indicators, and then calculate the comprehensive score of each industry. For the selection of target development industry nodes, those industries that perform well in the comprehensive score can be selected as the priority development direction. At the same time, considering the different resource endowments and development priorities of different regions, a personalized evaluation system can also be formulated for specific regions to better guide local industrial development.

[0083] Finally, taking the selected target development industry node as the end point and any major industry node in the region as the starting point, the path with the maximum complete consumption coefficient is selected as the prediction of the optimal industrial development path for the region. The "maximum complete consumption coefficient" here means that there is the strongest interdependence between the industries on the selected path, and therefore it is most likely to achieve coordinated development. By simulating the industrial development paths under different scenarios, a scientific basis can be provided for enterprises to plan market strategies and for investors to evaluate investment potential. Through the above methods and steps, a set of regional economic development indicator systems with both extensive coverage and in-depth analysis can be constructed, providing strong data support for the realization of accurate industrial development path predictions, and thus contributing to the healthy development of the regional economy.

[0084] In summary, the knowledge graph-based industrial development path prediction method based on the embodiment of the present application is explained, which first obtains the set of economic development indicators of the region to be analyzed, and introduces a data processing algorithm based on deep learning to perform semantic embedding coding, hub feature analysis and sparse processing on each economic development indicator, so as to dig out the overall characteristics of the regional economy and the law of industrial distribution, and then use the feedforward neural network model to set dynamic thresholds to screen out the main industries in the region, so as to predict the development path from the main industry to the target development industry on this basis. In this way, the diversity of economic development levels and industrial structures in different regions can be fully considered, and the industrial characteristics and development trends of each region can be dynamically adapted to improve the accuracy of industrial development path prediction.

[0085] Furthermore, a device for predicting industrial development paths based on a knowledge graph is provided, which can implement the method for predicting industrial development paths based on a knowledge graph as described above, and comprises: an economic development indicator acquisition module, used to acquire a set of regional economic development indicators; an economic development indicator embedding coding module, used to embed and code each regional economic development indicator in the set of regional economic development indicators to obtain a set of regional economic development indicator embedded coding vectors; a feature sparse processing module, used to perform sparse processing on the set of regional economic development indicator embedded coding vectors based on feature distribution hubs to obtain a set of sparse regional economic development indicator embedded coding vectors; a preset threshold setting module, used to set a preset threshold based on the set of sparse regional economic development indicator embedded coding vectors.

[0086] Here, those skilled in the art can understand that the specific operations of each module in the above-mentioned industrial development path prediction device based on knowledge graph have been referred to above. Figures 1 to 5 It has been introduced in detail in the description of the industrial development path prediction method based on knowledge graph, and therefore, its repeated description will be omitted.

[0087] The basic principle of the present invention is described above in conjunction with specific embodiments. However, it should be pointed out that the advantages, strengths, effects, etc. mentioned in the present invention are only examples and not limitations, and it cannot be considered that these advantages, strengths, effects, etc. must be possessed by each embodiment of the present invention. In addition, the specific details of the above embodiments are only for the purpose of illustration and facilitation of understanding, rather than limitation, and the above details do not limit the present invention to being implemented by adopting the above specific details.

[0088] In the above embodiments, the description of each embodiment has its own emphasis. For the parts that are not described or recorded in detail in a certain embodiment, please refer to the relevant description of other embodiments. In the several embodiments provided by the present invention, it should be understood that the disclosed system and method can be implemented in other ways. For example, the system embodiment described above is only schematic. For example, the unit division is only a logical function division, and there may be other division methods in actual implementation. The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.

[0089] It will be apparent to those skilled in the art that the invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description, and it is intended that all variations falling within the meaning and scope of the equivalent elements of the claims be included in the invention. Any reference to a figure in a claim should not be considered as limiting the claim to which it relates.

[0090] In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units stated in the system claims can also be implemented by one unit through software or hardware.

[0091] Finally, it should be noted that the above description has been given for the purpose of illustration and description. In addition, the above embodiments are only used to illustrate the technical solution of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, it should be understood by those skilled in the art that the technical solution of the present invention can be modified or replaced by equivalents without departing from the spirit and scope of the technical solution of the present invention.

Claims

1. A method for predicting industrial development paths based on knowledge graph, comprising: Collect massive amounts of industry and regional information data to build an industry and regional information database; An industrial knowledge graph is constructed based on the complete consumption coefficient between industries, and a regional knowledge graph is constructed based on the geographical proximity between regions; industries whose output value accounts for a greater than a predetermined threshold in a region are taken as the main industrial nodes of the region, and a regional industrial integration knowledge graph is constructed based on the output value proportion of the main industries in the region; the development scale index and development quality index of all industries are calculated to determine the target development industrial node of any region in the regional industrial integration knowledge graph, and the target development industrial node is taken as the end point, and any main industrial node of any region is taken as the starting point, and the path with the largest complete consumption coefficient is selected as the prediction of the optimal industrial development path of any region, characterized in that the setting of the preset threshold includes: Obtain a set of regional economic development indicators; Embedding and coding each regional economic development indicator in the regional economic development indicator set to obtain a set of regional economic development indicator embedded coding vectors; Performing a sparse processing on the set of regional economic development indicator embedded coding vectors based on feature distribution hubs to obtain a sparse set of regional economic development indicator embedded coding vectors; The preset threshold is set based on the set of sparsely embedded coding vectors of regional economic development indicators.

2. The method for predicting industrial development paths based on knowledge graph according to claim 1 is characterized in that: The set of regional economic development indicator embedded coding vectors is subjected to a sparse processing based on a feature distribution hub to obtain a sparse set of regional economic development indicator embedded coding vectors, including: Performing a feature distribution hub search on the set of regional economic development indicator embedding coding vectors to obtain a regional economic development indicator embedding distribution hub coding vector; Based on the feature offset of each regional economic development indicator embedded coding vector in the set of regional economic development indicator embedded coding vectors relative to the regional economic development indicator embedded distribution hub coding vector, the set of regional economic development indicator embedded coding vectors is subjected to feature sparse processing to obtain the sparse set of regional economic development indicator embedded coding vectors.

3. The method for predicting industrial development paths based on knowledge graph according to claim 2 is characterized in that: Performing a feature distribution hub search on the set of regional economic development indicator embedding coding vectors to obtain the regional economic development indicator embedding distribution hub coding vector, including: Calculating the characteristic difference factor of each regional economic development indicator embedded coding vector in the set of regional economic development indicator embedded coding vectors relative to the set of regional economic development indicator embedded coding vectors to obtain a set of regional economic development indicator semantic difference factors; Taking the inverse of each semantic difference factor of the regional economic development indicator in the set of semantic difference factors of the regional economic development indicator and performing normalization processing based on the Softmax function to obtain a set of semantic correlation factors of the hub features of the regional economic development indicator; Taking the set of semantic relevance factors of the hub features of the regional economic development indicators as the weight distribution, the set of embedded coding vectors of the regional economic development indicators is weightedly aggregated to obtain the embedded distribution hub coding vector of the regional economic development indicators.

4. The method for predicting industrial development paths based on knowledge graph according to claim 3 is characterized in that: Calculating the characteristic difference factor of each regional economic development indicator embedded coding vector in the set of regional economic development indicator embedded coding vectors relative to the set of regional economic development indicator embedded coding vectors to obtain a set of regional economic development indicator semantic difference factors, including: The mean of the Mahalanobis distances of each regional economic development indicator embedded coding vector in the set of regional economic development indicator embedded coding vectors relative to other regional economic development indicator embedded coding vectors is calculated respectively to obtain a set of semantic difference factors of the regional economic development indicators.

5. The method for predicting industrial development paths based on knowledge graph according to claim 4 is characterized in that: Based on the feature offset of each regional economic development indicator embedded coding vector in the set of regional economic development indicator embedded coding vectors relative to the regional economic development indicator embedded distribution hub coding vector, the set of regional economic development indicator embedded coding vectors is subjected to feature sparse processing to obtain the set of sparse regional economic development indicator embedded coding vectors, including: Respectively calculating the semantic fine-grained ablation factor of each regional economic development indicator embedding coding vector in the set of regional economic development indicator embedding coding vectors relative to the regional economic development indicator embedding distribution hub coding vector to obtain a set of regional economic development indicator semantic fine-grained ablation factors; Based on the set of semantic fine-grained ablation factors of the regional economic development indicators, feature sparse processing is performed on the set of regional economic development indicator embedded coding vectors to obtain the sparse set of regional economic development indicator embedded coding vectors.

6. The method for predicting industrial development paths based on knowledge graph according to claim 5 is characterized in that: Based on the set of semantic fine-grained ablation factors of the regional economic development indicators, feature sparse processing is performed on the set of regional economic development indicator embedded coding vectors to obtain the set of sparse regional economic development indicator embedded coding vectors, including: Inputting each of the regional economic development indicator semantic fine-grained ablation factors in the set of regional economic development indicator semantic fine-grained ablation factors into a sparse module based on a gating function to obtain a set of sparse regional economic development indicator semantic fine-grained ablation factors; Taking the set of the sparse regional economic development indicator semantic fine-grained ablation factors as the weight distribution, the set of the regional economic development indicator embedded coding vectors is weighted modulated to obtain the set of the sparse regional economic development indicator embedded coding vectors.

7. The method for predicting industrial development paths based on knowledge graph according to claim 6 is characterized in that: The preset threshold is set based on the set of sparsely embedded coding vectors of regional economic development indicators, including: The set of the sparsely-embedded coding vectors of regional economic development indicators is input into a threshold setter based on a feedforward neural network model to obtain the preset threshold.

8. An industrial development path prediction device based on knowledge graph, which can implement the industrial development path prediction method based on knowledge graph as described in any one of claims 1 to 7, characterized in that: The industrial development path prediction device based on knowledge graph includes: The economic development indicator acquisition module is used to obtain the regional economic development indicator set; An economic development indicator embedding coding module, used for embedding coding each regional economic development indicator in the regional economic development indicator set to obtain a set of regional economic development indicator embedding coding vectors; A feature sparse processing module, used for performing a sparse processing on the set of regional economic development indicator embedded coding vectors based on a feature distribution hub to obtain a sparse set of regional economic development indicator embedded coding vectors; The preset threshold setting module is used to set the preset threshold based on the set of sparsely embedded coding vectors of regional economic development indicators.

Citation Information

Patent Citations

  • Industrial development path prediction method and device based on knowledge graph

    CN112990575A

Cited By

  • New energy system development path prediction method, equipment and medium

    CN120996259A