Regional economic spatio-temporal knowledge graph index prediction method and device

By constructing a regional economic spatiotemporal knowledge graph and utilizing R-GCNs and XGBoost models, the problem of existing technologies being unable to capture the dynamics and multifaceted nature of economic activities has been solved, enabling dynamic analysis of the regional economy and accurate prediction of future indicators.

CN120875633APending Publication Date: 2025-10-31WUHAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510510361.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing regional economic models rely on static statistical indicators, such as GDP and employment rates, which fail to capture the dynamics and multifaceted nature of economic activity.

Method used

By employing a regional economic spatiotemporal knowledge graph, combined with relational graph convolutional networks (R-GCNs) and the XGBoost model, knowledge representation learning is performed through the construction of the spatiotemporal knowledge graph, cosine similarity is calculated, and targeted indicators for regional economic development are predicted.

Benefits of technology

It enables the dynamic and multifaceted capture of regional economic activities, provides more accurate forecasts of future trends and targeted indicator analysis, and supports policy making and resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120875633A_ABST
    Figure CN120875633A_ABST
Patent Text Reader

Abstract

The invention relates to an index prediction method and device for a regional economy space-time knowledge graph, and the method comprises the steps: constructing the regional economy space-time knowledge graph; performing knowledge representation learning on the spatial-temporal knowledge graph of the regional economy by using a preset relational graph convolutional network; determining target targeting indexes influencing regional economic development according to the cosine similarity; comparing index prediction degrees of a preset R-GCNs + XGBoost model, a preset random forest and a preset polynomial regression model for regional economic development; and predicting at least one key index of regional economy based on a preset R-GCNs + XGBoost model and the comparison data. According to the method, regional economy and the knowledge graph technology are creatively and deeply fused, inherent dynamic nature and space-time characteristics of economic development in a real scene are fully considered, and powerful technical support is provided for regional economy targeted index analysis and future index prediction by using powerful information organization and interpretation capabilities of the knowledge graph.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of geographic information science and technology, and in particular to a method and apparatus for predicting indicators of a regional economic spatiotemporal knowledge graph. Background Technology

[0002] Spatial-temporal Knowledge Graph (STKG) is an advanced data representation technique that integrates multi-source data and captures complex spatiotemporal relationships. Unlike traditional knowledge graphs that primarily focus on static entity relationships, STKG emphasizes the dynamic evolution of entities and their interactions over time and space. This capability makes them suitable for modeling phenomena exhibiting spatial heterogeneity and temporal dynamics, such as regional economic activities. Knowledge graph embedding is a core component of STKG, designed to embed entities and relationships into a low-dimensional vector space while preserving their semantic and structural information. This approach enables efficient reasoning, inference, and prediction tasks, which are crucial for analyzing complex systems such as regional economies. Currently, spatiotemporal knowledge graphs for regional economies are still lacking.

[0003] Regional economic analysis is a potential application area for spatiotemporal knowledge graphs (STKG). Traditional regional economic models typically rely on static statistical indicators such as GDP (Gross Domestic Product), industrial structure, and employment rates, which fail to capture the dynamic and multifaceted nature of economic activity. In contrast, STKG offers a comprehensive framework that integrates traditional statistical indicators with novel non-statistical indicators, such as nighttime light remote sensing, internet search indices, inter-city interactions, population flows, and extreme weather disturbances. This holistic approach based on "statistical + non-statistical" indicators allows for a more detailed analysis of regional economic phenomena, enabling researchers to uncover hidden patterns and more accurately predict future trends. Summary of the Invention

[0004] This invention provides a method and apparatus for predicting indicators of a regional economic spatiotemporal knowledge graph, in order to solve the problem that regional economic models in related technologies usually rely on static statistical indicators, such as GDP, industrial structure and employment rate, which cannot capture the dynamics and multifaceted nature of economic activities.

[0005] A first aspect of this invention provides a method for predicting indicators from a regional economic spatiotemporal knowledge graph, comprising the following steps: constructing a spatiotemporal knowledge graph of the regional economy based on entity determination information, relationship determination information, and attribute determination information of the regional economy; performing knowledge representation learning on the spatiotemporal knowledge graph of the regional economy using a preset relationship graph convolutional network to generate knowledge embedding features of entities and relationships; calculating the cosine similarity between economic entities and entities that meet preset other conditions based on the knowledge embedding features, and determining target indicators affecting the regional economic development based on the cosine similarity; comparing the predictive power of a preset XGBoost + R-GCNs model, a preset random forest, and a preset multinomial regression model for the indicators of the regional economic development based on the spatial differences and temporal variation patterns of the target indicators to generate comparative data; and predicting at least one key indicator of the regional economy based on the preset R-GCNs + XGBoost model and the comparative data.

[0006] Optionally, in one embodiment of the present invention, before constructing the spatiotemporal knowledge graph of the regional economy based on the determination of entities, relationships, and attributes of the regional economy, the method further includes: preprocessing at least one of the following multi-source heterogeneous regional economic data: statistical indicator data that meets preset traditional conditions, economic indicator data that meets preset new conditions, and target city spatial data, to generate the preprocessing result of the regional economy.

[0007] Optionally, in one embodiment of the present invention, constructing a spatiotemporal knowledge graph of the regional economy based on entity determination information, relationship determination information, and attribute determination information of the regional economy includes: determining entity determination information in the spatiotemporal knowledge graph of the regional economy according to spatial entities, temporal entities, and indicator entities; and constructing the spatiotemporal knowledge graph of the regional economy according to the entity determination information, the relationship determination information, and the attribute determination information.

[0008] Optionally, in one embodiment of the present invention, the step of using a preset relational graph convolutional network to perform knowledge representation learning on the regional economic spatiotemporal knowledge graph to generate knowledge embedding features of entities and relationships includes: using the aggregated time information, spatial information, and indicator information of the regional economy to characterize the dynamic features of the regional economy; based on the dynamic features, using the preset relational graph convolutional network to perform knowledge representation learning on the regional economic spatiotemporal knowledge graph to generate learning results; and based on the learning results, using a preset message passing mechanism to capture the high-order relationships and spatiotemporal dependencies between entities to generate knowledge embedding features of the entities and relationships.

[0009] Optionally, in one embodiment of the present invention, the calculation formula for the knowledge embedding feature is:

[0010] in, For relationship specific Weight matrix, Indicates through the relationship Connected nodes The neighborhood group, As the normalization factor, For activation function, For the set of all relations, For the target node and its neighboring nodes, Let be the representation of node u at layer l.

[0011] Optionally, in one embodiment of the present invention, the formula for calculating the cosine similarity is:

[0012] in, and These are the embedding vectors for the entity and other types of entities, respectively.

[0013] Optionally, in one embodiment of the present invention, the influence factor of the preset XGBoost model + R-GCNs model is:

[0014] in, For the predicted value of the sample, The input features of the sample are... The target indicator is... Let be the prediction function of the k-th decision tree.

[0015] A second aspect of the present invention provides an indicator prediction device for a regional economic spatiotemporal knowledge graph, comprising: a construction module for constructing a spatiotemporal knowledge graph of the regional economy based on entity determination information, relationship determination information, and attribute determination information of the regional economy; a generation module for performing knowledge representation learning on the spatiotemporal knowledge graph of the regional economy using a preset relationship graph convolutional network to generate knowledge embedding features of entities and relationships; a calculation module for calculating the cosine similarity between economic entities and entities that meet preset other conditions based on the knowledge embedding features, and determining target indicators affecting the regional economic development based on the cosine similarity; a comparison module for comparing the predictive power of a preset XGBoost model + R-GCNs model, a preset random forest model, and a preset multinomial regression model for the indicators of the regional economic development based on the spatial differences and temporal variation patterns of the target indicators to generate comparison data; and a prediction module for predicting at least one key indicator of the regional economy based on the preset R-GCNs + XGBoost model and the comparison data.

[0016] Optionally, in one embodiment of the present invention, it further includes: a preprocessing module, used to preprocess at least one of the following multi-source heterogeneous regional economic data, namely, statistical indicator data that meets preset traditional conditions, economic indicator data that meets preset new conditions, and target city spatial data, before constructing the spatiotemporal knowledge graph of the regional economy based on entity determination, relationship determination, and attribute determination of the regional economy, so as to generate the preprocessing result of the regional economy.

[0017] Optionally, in one embodiment of the present invention, the construction module includes: a determining unit, configured to determine entity determination information in the regional economic spatiotemporal knowledge graph based on spatial entities, temporal entities, and indicator entities; and a construction unit, configured to construct the regional economic spatiotemporal knowledge graph based on the entity determination information, the relationship determination information, and the attribute determination information.

[0018] Optionally, in one embodiment of the present invention, the resource generation module includes: a characterization unit, used to characterize the dynamic features of the regional economy using aggregated time information, spatial information, and indicator information of the regional economy; a learning unit, used to perform knowledge representation learning on the spatiotemporal knowledge graph of the regional economy based on the dynamic features using the preset relationship graph convolutional network, so as to generate learning results; and a generation unit, used to capture the high-order relationships and spatiotemporal dependencies between the entities based on the learning results using a preset message passing mechanism, so as to generate knowledge embedding features of the entities and the relationships.

[0019] Optionally, in one embodiment of the present invention, the calculation formula for the knowledge embedding feature is:

[0020] in, For relationship specific Weight matrix, Indicates through the relationship Connected nodes The neighborhood group, As the normalization factor, For activation function, For the set of all relations, For the target node and its neighboring nodes, Let be the representation of node u at layer l.

[0021] Optionally, in one embodiment of the present invention, the formula for calculating the cosine similarity is:

[0022] in, and These are the embedding vectors for the entity and other types of entities, respectively.

[0023] Optionally, in one embodiment of the present invention, the influence factor of the preset XGBoost model + R-GCNs model is:

[0024] in, For the predicted value of the sample, The input features of the sample are... The target indicator is... Let be the prediction function of the k-th decision tree.

[0025] A third aspect of the present invention provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the regional economic spatiotemporal knowledge graph indicator prediction method as described in the above embodiments.

[0026] A fourth aspect of the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for predicting indicators of a regional economic spatiotemporal knowledge graph.

[0027] This invention innovatively integrates regional economics with knowledge graph technology, fully considering the inherent dynamics and spatiotemporal characteristics of economic development in real-world scenarios. Leveraging the powerful information organization and interpretation capabilities of knowledge graphs, it provides strong technical support for targeted indicator analysis and future indicator prediction in regional economics. This solves the problem that related regional economic models often rely on static statistical indicators, such as GDP, industrial structure, and employment rate, which fail to capture the dynamics and multifaceted nature of economic activities.

[0028] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0029] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 A flowchart illustrating an indicator prediction method for a regional economic spatiotemporal knowledge graph according to an embodiment of the present invention; Figure 2 This is a flowchart illustrating the overall process of a regional economic spatiotemporal knowledge graph indicator prediction method according to an embodiment of the present invention. Figure 3 A regional economic knowledge graph pattern layer diagram for a regional economic spatiotemporal knowledge graph index prediction method according to an embodiment of the present invention. Figure 4 This diagram illustrates the selection of statistical and non-statistical indicators for an indicator prediction method of a regional economic spatiotemporal knowledge graph according to an embodiment of the present invention. Figure 5 A flowchart of knowledge representation learning (graph convolutional neural network) for an indicator prediction method of regional economic spatiotemporal knowledge graph according to an embodiment of the present invention; Figure 6 This is a diagram showing the target indicator identification results of a regional economic spatiotemporal knowledge graph indicator prediction method according to an embodiment of the present invention. Figure 7 This is a graph showing the time variation of a target indicator in a method for predicting indicators of a regional economic spatiotemporal knowledge graph according to an embodiment of the present invention. Figure 8 This is a comparison diagram of the model results of the regional economic spatiotemporal knowledge graph index prediction method according to an embodiment of the present invention. Figure 9 This is a schematic diagram of the structure of an indicator prediction device for a regional economic spatiotemporal knowledge graph according to an embodiment of the present invention; Figure 10 This is a schematic diagram of the structure of an electronic device provided according to an embodiment of the present invention. Detailed Implementation

[0030] Embodiments of the present invention are described in detail below. Examples of these embodiments are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0031] The following description, with reference to the accompanying drawings, illustrates a method and apparatus for predicting indicators using a regional economic spatiotemporal knowledge graph, as described in this invention. Addressing the issue that regional economic models in related technologies, as mentioned in the background section, typically rely on static statistical indicators such as GDP, industrial structure, and employment rate, which fail to capture the dynamics and multifaceted nature of economic activities, this invention provides a method for predicting indicators using a regional economic spatiotemporal knowledge graph. This method innovatively integrates regional economics with knowledge graph technology, fully considering the inherent dynamics and spatiotemporal characteristics of economic development in real-world scenarios. Utilizing the powerful information organization and interpretation capabilities of knowledge graphs, it provides strong technical support for targeted indicator analysis and future indicator prediction in the regional economy. This solves the problem that regional economic models in related technologies typically rely on static statistical indicators such as GDP, industrial structure, and employment rate, which fail to capture the dynamics and multifaceted nature of economic activities.

[0032] Specifically, Figure 1 This is a flowchart illustrating a method for predicting indicators of a regional economic spatiotemporal knowledge graph, as provided in an embodiment of the present invention.

[0033] like Figure 1 As shown, the indicator prediction method of the regional economic spatiotemporal knowledge graph includes the following steps: In step S101, a spatiotemporal knowledge graph of the regional economy is constructed based on the entity determination information, relationship determination information, and attribute determination information of the regional economy.

[0034] In actual implementation, such as Figure 2As shown, this embodiment of the invention can construct a regional economic spatiotemporal knowledge graph pattern layer and a data layer. Entity determination, relationship determination, and attribute determination information for the regional economy are obtained through steps such as entity determination, relationship determination, and attribute determination. Based on this information, a spatiotemporal knowledge graph of the regional economy is constructed. The data layer of the regional economic spatiotemporal knowledge graph is constructed based on datasets from the *China Statistical Yearbook*, *China County Statistical Yearbook*, NPP-VIIRS (National Polar-orbiting Partnership-Visible Infrared Imaging Radiometer) nighttime light dataset, and extreme climate datasets. The data is stored in the form of triples, such as... Figure 3 As shown.

[0035] Optionally, in one embodiment of the present invention, before constructing the spatiotemporal knowledge graph of the regional economy based on the determination of entities, relationships, and attributes of the regional economy, the method further includes: preprocessing at least one of the multi-source heterogeneous regional economic data among statistical indicator data that meets preset traditional conditions, economic indicator data that meets preset new conditions, and target city spatial data, to generate a preprocessing result of the regional economy.

[0036] It is understood that, in the embodiments of the present invention, statistical indicator data that meet the preset traditional conditions can be traditional statistical indicator data; economic indicator data that meet the preset new conditions can be new economic indicator data; and the target city spatial data can be Chinese city spatial data. The present invention uses Chinese city-level cities as the analysis unit, and first constructs a regional economic spatiotemporal knowledge graph based on long-term series traditional statistical indicators and new non-statistical indicators, as well as the spatial accessibility relationships between cities.

[0037] Specifically, such as Figure 4 As shown, embodiments of the present invention can acquire multi-source heterogeneous data and preprocess multi-source heterogeneous regional economic data (traditional statistical indicator data, new economic indicator data, and spatial data of Chinese cities) using both statistical and non-statistical methods. This includes operations such as data cleaning and normalization to generate preprocessed results of regional economic data, laying the foundation for the subsequent construction of a spatiotemporal knowledge graph.

[0038] Blank values ​​in statistical data are filled using regression imputation. The specific calculation method is as follows:

[0039] in, yes Predicted values ​​(used to fill in missing values); It is the intercept term; , It is the regression coefficient; , ,…, It is the independent variable; It is a random error term, and its mean is usually assumed to be 0.

[0040] The total and mean values ​​of nighttime light data and extreme climate data were statistically analyzed by region. The extreme climate indices selected were CDD and FD0. CDD represents the longest duration of the dry season, and FD0 represents the longest duration of frost, specifically the number of days with a daily minimum temperature (TN) below 0 degrees Celsius. The formulas are as follows:

[0041]

[0042] in, For the total statistical value, This is the statistical mean; It is a partition The grid, yes Total number of inner grid cells For the row number of the raster cell, For the grid cell column number, Indicates which partition.

[0043] This invention uses a regression imputation method to fill in blank values ​​in statistical data, ensuring data integrity. For nighttime light data and extreme climate data, regional statistics are performed to calculate the total and mean values. The selected extreme climate indicators include CDD (Consecutive Dry Days, the longest duration of a dry period) and FD0 (Frost Days, the longest duration of frost).

[0044] Optionally, in one embodiment of the present invention, a spatiotemporal knowledge graph of the regional economy is constructed based on entity determination information, relationship determination information, and attribute determination information of the regional economy, including: determining entity determination information in the spatiotemporal knowledge graph of the regional economy according to spatial entities, temporal entities, and indicator entities; and constructing the spatiotemporal knowledge graph of the regional economy according to entity determination information, relationship determination information, and attribute determination information.

[0045] In actual implementation, embodiments of the present invention can determine entity identification information in the regional economic spatiotemporal knowledge graph based on spatial entities, temporal entities, and indicator entities, and construct the regional economic spatiotemporal knowledge graph based on entity identification information, relationship identification information, and attribute identification information.

[0046] The method for constructing the spatiotemporal knowledge graph schema layer for regional economies mainly includes identifying indicator entities, spatial entities, and temporal entities, and determining the attribute values ​​for each entity. Secondly, it defines the types of relationships between entities, including relationships between temporal entities and spatial entities, and relationships between spatial entities. Actual regional economic data is imported and stored in triples based on the schema layer. The specific formula is as follows:

[0047]

[0048] in, For a collection of entities, For a set of relations, For a collection of attributes, A triplet of facts.

[0049] In step S102, a pre-defined relational graph convolutional network is used to learn the knowledge representation of the spatiotemporal knowledge graph of the regional economy in order to generate knowledge embedding features of entities and relationships.

[0050] It is understood that the preset relational graph convolutional network in this embodiment of the invention can be R-GCNs. Entities in the regional economic spatiotemporal knowledge graph are divided into three categories: spatial entities (e.g., cities), temporal entities, and indicator entities (e.g., GDP entities and population size entities). Each entity is associated with specific attributes, such as the attributes of an investment entity, including real estate development investment, total social fixed asset investment, and completed urban fixed asset investment. Relationships in the spatiotemporal knowledge graph include spatial relationships, temporal relationships, and indicator relationships. These entities, relationships, and their attributes together constitute the basic framework of the regional economic spatiotemporal knowledge graph.

[0051] In practical implementation, this invention can perform knowledge representation learning for regional economic spatiotemporal knowledge graphs. After the knowledge graph is constructed, R-GCNs (Relational Graph Convolutional Networks) are used to learn the knowledge representation of the regional economic spatiotemporal knowledge graph, obtaining knowledge embedding features of entities and relationships. R-GCNs extends the traditional GCN (Graph Convolutional Network) by merging relationship-specific weight matrices to handle heterogeneous edge types, making it suitable for modeling knowledge graphs and relational data. The specific steps include designing a message passing mechanism, selecting graph neural network hyperparameters such as the number of neighbors in the first and second layers (in this invention, the number of neighbors in the first layer is 5, and in the second layer, it is 3), selecting a suitable loss function (in this invention, negative sampling loss is selected), and finally designing a reasonable number of network layers.

[0052] Optionally, in one embodiment of the present invention, a pre-defined relational graph convolutional network is used to perform knowledge representation learning on the spatiotemporal knowledge graph of the regional economy to generate knowledge embedding features of entities and relationships. This includes: using the aggregation time information, spatial information, and indicator information of the regional economy to characterize the dynamic features of the regional economy; based on the dynamic features, using the pre-defined relational graph convolutional network to perform knowledge representation learning on the spatiotemporal knowledge graph of the regional economy to generate learning results; and based on the learning results, using a pre-defined message passing mechanism to capture high-order relationships and spatiotemporal dependencies between entities to generate knowledge embedding features of entities and relationships.

[0053] In practical implementation, embodiments of this invention can comprehensively characterize the dynamic features of a regional economy by aggregating temporal information (such as time entities and attributes), spatial information (such as geographical regions and relationships), and indicator information (such as economic indicators and attributes) when constructing a regional economic spatiotemporal knowledge graph. Knowledge representation learning is performed using Relational Graph Convolutional Networks (R-GCNs), and a message passing mechanism is used to effectively capture high-order relationships and spatiotemporal dependencies between entities, generating high-quality knowledge embeddings, such as... Figure 5 As shown, in unsupervised training, dynamic negative sampling and relation-aware negative sampling methods are innovatively adopted to optimize model performance. Simultaneously, by reasonably selecting neighborhood hyperparameters (such as choosing the neighborhood size (5, 3)), the generalization ability and prediction accuracy of the model are further improved, providing strong support for the identification of regional economic target indicators and future prediction. The knowledge embeddings of each indicator node and city node generated by R-GCNs can be used as input features for further analysis in subsequent knowledge reasoning.

[0054] In one embodiment of the present invention, an unsupervised relational graph convolutional network (R-GCNs) with a negative sampling loss function is employed. R-GCNs extend traditional GCNs to handle graphs with multiple edge types, making them well-suited for modeling knowledge graphs and relational data. Unlike standard GCNs that assume homogeneous graphs, R-GCNs handle heterogeneous edge types by incorporating relation-specific weight matrices, thereby learning the influence of various relations on node representations. Mathematically, the feature results of the knowledge embedding of each indicator entity and spatiotemporal entity can be calculated. The formula for calculating the knowledge embedding features is:

[0055] in, For relationship specific Weight matrix, Indicates through the relationship Connected nodes The neighborhood group, As the normalization factor, For activation function, For the set of all relations, For the target node and its neighboring nodes, Let be the representation of node u at layer l.

[0056] It should be noted that R-GCNs update the feature representation of nodes through a message passing mechanism. Its core idea is to aggregate information from neighboring nodes and update it in combination with the relationship type.

[0057] R-GCNs typically use a negative sampling loss function to train the model, aiming to maximize the score of positive samples while minimizing the score of negative samples. The formula for the negative sampling loss function is as follows:

[0058] In the formula, The set of positive samples represents the truly existing triples. ), For head entity, For relationships, It is a tail entity; Let be the negative sample set, representing randomly generated pseudo triples (( ', , ') ); Let the scoring function of the triple be the first term. This represents maximizing the score of the positive samples; the second term This represents minimizing the score of negative samples; achieved by optimizing the loss function. The model is able to learn effective embedded representations of entities and relationships.

[0059] In step S103, based on knowledge embedding features, the cosine similarity between the economic entity and entities that meet the remaining preset conditions is calculated, and the target indicators that affect regional economic development are determined based on the cosine similarity.

[0060] It is understood that the entities for which other conditions are preset in the embodiments of the present invention can be entities other than economic entities.

[0061] As one possible implementation method, embodiments of the present invention can perform targeted indicator identification. Based on knowledge embedding features, knowledge reasoning is performed to calculate the cosine similarity between economic entities and other entities that meet certain criteria (such as population size entities, investment entities, etc.). Simultaneously, the average cosine similarity of each type of entity over the research period is ranked, and the main targeted indicators affecting regional economic development are calculated and revealed through similarity analysis.

[0062] It should be noted that the other preset conditions can be set by those skilled in the art according to the actual situation, and no specific restrictions are imposed here.

[0063] In one embodiment of the invention, cosine similarity is used to quantify the relationship strength between the target entity and other types of entities, and the target indicators are screened based on the relationship strength. The PDF (Probability Density Function) of a normal distribution is calculated based on the cosine similarity of all entities to obtain the mean and variance of the cosine similarity. A threshold is set to the sum of the mean and variance to obtain the proportion of each type of entity whose cosine similarity to the economic entity (target entity) is greater than the threshold. This is used to determine the target indicators for economic growth. By setting the threshold, target indicators that have a significant impact on economic growth are screened out, such as... Figure 6 As shown, the formula for calculating the cosine similarity between two vectors is:

[0064] in, and These are the embedding vectors for the entity and other types of entities, respectively. The dot product measures the similarity between the vectors, and the denominator normalizes the vectors to ensure that the cosine similarity is within the range of (-1, 1).

[0065] The specific formula for the probability density function of the normal distribution is:

[0066] in, It is a random variable; It is the mean (the central location of the distribution); It is the standard deviation (a measure of the dispersion of data); π is the variance; π is the mathematical constant pi (approximately 3.14159); exp is the exponential function.

[0067] In step S104, based on the spatial differences and temporal variation patterns of the target indicators, the predictive power of the preset XGBoost model + R-GCNs model, the preset random forest model, and the preset multinomial regression model on regional economic development indicators is compared to generate comparative data.

[0068] It is understood that XGBoost in this embodiment of the invention is an ensemble learning method based on the gradient boosting algorithm, suitable for regression and classification tasks; Random Forest makes predictions by constructing multiple decision trees, and has anti-overfitting capabilities; Multinomial Regression captures nonlinear relationships in the data by introducing multinomial features. Through comparison, the optimal model is selected for predicting future economic indicators. The detailed steps are as follows.

[0069] 1) Dataset Preparation: A sliding window method is used to divide the data into multiple training and test sets by sliding a fixed-length window across the time series. Each window contains a continuous segment of historical data as the training set, and the subsequent future data as the test set. This invention uses data from the past 5 years to predict data from the next 3 years.

[0070] 2) Model Training: XGBoost (eXtreme Gradient Boosting), Random Forest (RF), and Multinomial Regression were selected respectively. Forward Propagation: Input node embeddings were performed, and output probabilities were calculated. Backward Propagation: The gradient was calculated based on the loss function, and model parameters were updated. Validation: Model performance was monitored using a validation set, and hyperparameters (such as learning rate, number of hidden layers, etc.) were adjusted. Taking XGBoost as an example, XGBoost is an ensemble learning method based on the gradient boosting algorithm, suitable for regression and classification tasks. It iteratively trains multiple decision trees, gradually reducing the model's error. The goal of XGBoost is to minimize the loss function, which includes prediction loss and regularization terms.

[0071]

[0072] Where K is the total number of trees; Which tree is it? For the i-th sample, It is the actual value. It is a predicted value. For loss, It is the loss function (such as mean squared error), where n is the number of samples. It is the j-th tree. It is a regularization term used to control model complexity.

[0073] Secondly, there's the tree construction. In each iteration, XGBoost adds a new tree by minimizing the objective function. The tree gain can be expressed as:

[0074] Where G is the gradient (the derivative of the loss function with respect to the predicted value), and H is the Hessian matrix (the second derivative of the loss function). It is the regularization parameter.

[0075] 3) Performance Evaluation. Data from 2019-2021 is used as the test set to evaluate model performance. Common evaluation metrics include: Mean Squared Error (MSE) and Mean Absolute Percentage Error (MAPE). MSE is the average of the squared differences between predicted and true values, used to measure the accuracy of the prediction model. MAPE is the average of the absolute percentage differences between predicted and true values, used to measure the relative error of the prediction model.

[0076]

[0077]

[0078] in, It is the true value of the i-th sample; is the predicted value of the i-th sample; N is the number of samples.

[0079] Random Forest (RF) is an ensemble learning method that performs classification or regression by constructing multiple decision trees and combining their outputs. It exhibits resistance to overfitting and high prediction accuracy. Random Forest uses "out-of-bag samples" and a voting mechanism for prediction. For regression tasks, the output is the average of the predictions from all the trees.

[0080] in, It is the number of trees. It is the predicted value of the t-th tree.

[0081] Next, information gain is applied. When constructing each tree, the random forest uses information gain or the Gini index to select the optimal splitting feature. Information gain can be expressed as:

[0082] in, It is a dataset entropy, It is a characteristic. It is a feature Values A subset of.

[0083] Polynomial linear regression is an extended linear regression method that captures nonlinear relationships in data by introducing polynomial features. It is suitable for scenarios where data exhibits complex nonlinear trends, and can fit more complex curves through polynomial terms. The goal of polynomial linear regression is to minimize the sum of squared residuals. The model of polynomial linear regression can be expressed as:

[0084] Where y is the dependent variable and x is the independent variable; , It is the regression coefficient; , ,…, It is the independent variable; It is the random error term; RSS is the minimum sum of squared residuals. It is the actual value; It is a predicted value; This refers to the sample size. The regression coefficients are solved using the Ordinary Least Squares (OLS) method.

[0085] In practical implementation, embodiments of the present invention can perform spatial differences and temporal variations analysis of targeted indicators. For example... Figure 7 As shown, this study delves into the spatial differences and temporal variation patterns of target indicators to provide a scientific basis for coordinated regional economic development. Based on these patterns, it compares the accuracy of XGBoost + R-GCNs model, Random Forest (RF), and Multinomial Regression model in predicting future indicators of regional economic development. Comparative data is generated to verify the reliability of R-GCNs + XGBoost. Figure 8 As shown.

[0086] In one embodiment of the present invention, R-GCNs+XGBoost is used to predict future economic indicators. The node embeddings generated by R-GCNs and the target indicators obtained from the above steps are used as input features of XGBoost. That is, the preset influence factor of the XGBoost model + R-GCNs model is:

[0087] in, For the predicted value of the sample, For the input features of the sample, For target indicators, Let be the prediction function of the k-th decision tree.

[0088] In step S105, based on the preset R-GCNs+XGBoost model and comparative data, at least one key indicator of the regional economy is predicted.

[0089] Specifically, embodiments of the present invention can predict key regional economic indicators (such as GDP and GDP growth rate) based on an R-GCNs+XGBoost model and comparative data. Using 2021 data as input, the trained R-GCNs+XGBoost model outputs a GDP forecast for the next three years. This model combines the semantic information of knowledge graphs with the efficient predictive capabilities of XGBoost, providing accurate economic indicator forecasts and valuable decision support for policymakers and stakeholders.

[0090] This invention proposes a novel spatiotemporal knowledge graph model specifically tailored for regional economic analysis. The model not only represents the complex relationships between economic entities but also utilizes knowledge graph representation learning to predict potential relationships and future economic indicators. For example, it can infer the impact of policy changes on regional economic growth or predict the evolution of key economic indicators such as GDP and employment rates. By combining qualitative and quantitative predictive capabilities, the model can provide valuable insights for policymakers and stakeholders, promoting informed decision-making and resource allocation.

[0091] The application of STKG in regional economic analysis is particularly important in the current context due to the significant differences between regions. For example, different regions exhibit different economic roles and development patterns, and the model of this invention helps clarify these differences by analyzing development paths and identifying key factors driving economic growth. Furthermore, it provides a forward-looking perspective by predicting future economic types and trends, such as the emergence of regional economic centers or innovation-driven cities.

[0092] The regional economic spatiotemporal knowledge graph-based indicator prediction method proposed in this invention innovatively integrates regional economics with knowledge graph technology. It fully considers the inherent dynamics and spatiotemporal characteristics of economic development in real-world scenarios, and leverages the powerful information organization and interpretation capabilities of knowledge graphs to provide strong technical support for targeted indicator analysis and future indicator prediction in regional economies. This solves the problem that regional economic models in related technologies typically rely on static statistical indicators, such as GDP, industrial structure, and employment rate, which cannot capture the dynamics and multifaceted nature of economic activities.

[0093] Next, referring to the accompanying drawings, we describe the indicator prediction device for the regional economic spatiotemporal knowledge graph proposed according to an embodiment of the present invention.

[0094] Figure 9 This is a schematic diagram of the structure of the regional economic spatiotemporal knowledge graph indicator prediction device according to an embodiment of the present invention.

[0095] like Figure 9As shown, the indicator prediction device 10 of the regional economic spatiotemporal knowledge graph includes: a construction module 100, a generation module 200, a calculation module 300, a comparison module 400, and a prediction module 500.

[0096] Specifically, module 100 is used to construct a spatiotemporal knowledge graph of the regional economy based on entity determination information, relationship determination information, and attribute determination information of the regional economy.

[0097] The generation module 200 is used to learn the knowledge representation of the spatiotemporal knowledge graph of the regional economy using a pre-defined relational graph convolutional network, so as to generate knowledge embedding features of entities and relationships.

[0098] The calculation module 300 is used to calculate the cosine similarity between an economic entity and an entity that meets other preset conditions based on knowledge embedding features, and to determine the target indicators that affect regional economic development based on the cosine similarity.

[0099] The comparison module 400 is used to compare the predictive power of preset XGBoost + R-GCNs model, preset random forest and preset multinomial regression model on regional economic development indicators based on the spatial differences and temporal variation patterns of target indicators, so as to generate comparative data.

[0100] The prediction module 500 is used to predict at least one key indicator of the regional economy based on a preset R-GCNs+XGBoost model and comparative data.

[0101] Optionally, in one embodiment of the present invention, the indicator prediction device 10 for regional economic spatiotemporal knowledge graph further includes a preprocessing module.

[0102] The preprocessing module is used to preprocess at least one of the following multi-source heterogeneous regional economic data before constructing a spatiotemporal knowledge graph of the regional economy based on the determination of entities, relationships, and attributes of the regional economy: statistical indicator data that meets preset traditional conditions, economic indicator data that meets preset new conditions, and spatial data of the target city area, so as to generate the preprocessing result of the regional economy.

[0103] Optionally, in one embodiment of the present invention, the construction module 100 includes a determining unit and a construction unit.

[0104] Among them, the determination unit is used to determine the entity determination information in the regional economic spatiotemporal knowledge graph based on spatial entities, temporal entities, and indicator entities.

[0105] The building unit is used to construct a spatiotemporal knowledge graph of the regional economy based on entity-specific information, relationship-specific information, and attribute-specific information.

[0106] Optionally, in one embodiment of the present invention, the generation module 200 includes: a characterization unit, a learning unit, and a generation unit.

[0107] Among them, the characterization unit is used to characterize the dynamic features of the regional economy by utilizing aggregated time information, spatial information and indicator information of the regional economy.

[0108] The learning unit is used to learn knowledge representations of regional economic spatiotemporal knowledge graphs based on dynamic features and using a pre-defined relational graph convolutional network to generate learning results.

[0109] The generation unit is used to capture high-order relationships and spatiotemporal dependencies between entities based on the learning results and using a preset message passing mechanism to generate knowledge embedding features of entities and relationships.

[0110] Optionally, in one embodiment of the present invention, the formula for calculating the knowledge embedding feature is:

[0111] in, For relationship specific Weight matrix, Indicates through the relationship Connected nodes The neighborhood group, As the normalization factor, For activation function, For the set of all relations, For the target node and its neighboring nodes, Let be the representation of node u at layer l.

[0112] Optionally, in one embodiment of the present invention, the formula for calculating cosine similarity is:

[0113] in, and These are the embedding vectors for entities and other types of entities, respectively.

[0114] Optionally, in one embodiment of the present invention, the preset influence factor of the XGBoost model + R-GCNs model is:

[0115] in, For the predicted value of the sample, For the input features of the sample, For target indicators, Let be the prediction function of the k-th decision tree.

[0116] It should be noted that the explanation of the aforementioned embodiment of the indicator prediction method for regional economic spatiotemporal knowledge graph also applies to the indicator prediction device for regional economic spatiotemporal knowledge graph in this embodiment, and will not be repeated here.

[0117] The regional economic spatiotemporal knowledge graph indicator prediction device proposed in this invention innovatively integrates regional economics with knowledge graph technology. It fully considers the inherent dynamics and spatiotemporal characteristics of economic development in real-world scenarios, and leverages the powerful information organization and interpretation capabilities of knowledge graphs to provide strong technical support for targeted indicator analysis and future indicator prediction in regional economies. This solves the problem that regional economic models in related technologies typically rely on static statistical indicators, such as GDP, industrial structure, and employment rate, which cannot capture the dynamics and multifaceted nature of economic activities.

[0118] Figure 10 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. The electronic device may include: The memory 1001, the processor 1002, and the computer program stored on the memory 1001 and capable of running on the processor 1002.

[0119] When the processor 1002 executes the program, it implements the regional economic spatiotemporal knowledge graph indicator prediction method provided in the above embodiments.

[0120] Furthermore, electronic devices also include: Communication interface 1003 is used for communication between memory 1001 and processor 1002.

[0121] The memory 1001 is used to store computer programs that can run on the processor 1002.

[0122] The memory 1001 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0123] If the memory 1001, processor 1002, and communication interface 1003 are implemented independently, then the communication interface 1003, memory 1001, and processor 1002 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized into address buses, data buses, control buses, etc. For ease of representation, Figure 10 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0124] Optionally, in a specific implementation, if the memory 1001, processor 1002, and communication interface 1003 are integrated on a single chip, then the memory 1001, processor 1002, and communication interface 1003 can communicate with each other through an internal interface.

[0125] The processor 1002 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement embodiments of the present invention.

[0126] This embodiment also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for predicting indicators of a regional economic spatiotemporal knowledge graph.

[0127] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0128] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0129] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.

[0130] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0131] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any of the following techniques known in the art, or a combination thereof: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0132] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0133] Furthermore, the functional units in the various embodiments of the present invention can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0134] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for predicting indicators using a regional economic spatiotemporal knowledge graph, characterized in that, Includes the following steps: Based on the entity identification information, relationship identification information, and attribute identification information of the regional economy, a spatiotemporal knowledge graph of the regional economy is constructed. A pre-defined relational graph convolutional network is used to learn knowledge representations of the spatiotemporal knowledge graph of the regional economy, so as to generate knowledge embedding features of entities and relationships. Based on the knowledge embedding features, the cosine similarity between the economic entity and entities that meet the remaining preset conditions is calculated, and the target indicators that affect the regional economic development are determined based on the cosine similarity. Based on the spatial differences and temporal variation patterns of the target indicators, the predictive power of the preset XGBoost + R-GCNs model, preset random forest, and preset multinomial regression model on the indicators of regional economic development is compared to generate comparative data. Based on the preset R-GCNs+XGBoost model and the comparison data, at least one key indicator of the regional economy is predicted.

2. The indicator prediction method for regional economic spatiotemporal knowledge graphs according to claim 1, characterized in that, Before constructing the spatiotemporal knowledge graph of the regional economy based on entity identification, relationship identification, and attribute identification, the following steps are also included: At least one of the following multi-source heterogeneous regional economic data—statistical indicator data that meets preset traditional conditions, economic indicator data that meets preset new conditions, and spatial data of the target city area—is preprocessed to generate the preprocessed result of the regional economy.

3. The indicator prediction method for regional economic spatiotemporal knowledge graphs according to claim 1, characterized in that, The aforementioned construction of a spatiotemporal knowledge graph of the regional economy based on entity identification information, relationship identification information, and attribute identification information includes: The entity identification information in the regional economic spatiotemporal knowledge graph is determined based on spatial entities, temporal entities, and indicator entities. Construct a spatiotemporal knowledge graph of the regional economy based on the entity determination information, the relationship determination information, and the attribute determination information.

4. The indicator prediction method for regional economic spatiotemporal knowledge graphs according to claim 1, characterized in that, The step of using a pre-defined relational graph convolutional network to learn knowledge representations for the regional economic spatiotemporal knowledge graph, in order to generate knowledge embedding features of entities and relationships, includes: By utilizing the aggregated temporal, spatial, and indicator information of the regional economy, the dynamic characteristics of the regional economy are characterized. Based on the dynamic features, the preset relational graph convolutional network is used to learn the knowledge representation of the regional economic spatiotemporal knowledge graph to generate learning results; Based on the learning results, a preset message passing mechanism is used to capture the higher-order relationships and spatiotemporal dependencies between the entities, so as to generate knowledge embedding features of the entities and the relationships.

5. The indicator prediction method for regional economic spatiotemporal knowledge graphs according to claim 1, characterized in that, The formula for calculating the knowledge embedding feature is as follows: Among them, W r For the weight matrix specific to relation r, This represents the set of neighbors of node v connected through the relation r. Let be the normalization factor, σ be the activation function, R be the set of all relations, and u be the target node and its neighboring nodes. Let be the representation of node u at layer l.

6. The indicator prediction method for regional economic spatiotemporal knowledge graphs according to claim 1, characterized in that, The formula for calculating the cosine similarity is: Among them, H v-target and H v-other These are the embedding vectors for the entity and other types of entities, respectively.

7. The indicator prediction method for regional economic spatiotemporal knowledge graphs according to claim 1, characterized in that, The influence factor of the preset XGBoost module + R-GCNs model is: Among them, y predict Let x be the predicted value of the sample, and target1, target2...target2 be the input feature of the sample. n f is the target indicator. k Let be the prediction function of the k-th decision tree.

8. A predictive device for indicators of a regional economic spatiotemporal knowledge graph, characterized in that, include: The construction module is used to construct a spatiotemporal knowledge graph of the regional economy based on entity identification information, relationship identification information, and attribute identification information of the regional economy. The generation module is used to learn the knowledge representation of the spatiotemporal knowledge graph of the regional economy using a preset relation graph convolutional network, so as to generate knowledge embedding features of entities and relationships. The calculation module is used to calculate the cosine similarity between the economic entity and the entity that meets the other preset conditions based on the knowledge embedding features, and to determine the target indicators that affect the economic development of the region based on the cosine similarity. The comparison module is used to compare the predictive power of the preset XGBoost model + R-GCNs model, preset random forest and preset multinomial regression model on the indicators of regional economic development based on the spatial differences and temporal variation patterns of the target indicators, so as to generate comparison data. The prediction module is used to predict at least one key indicator of the regional economy based on a preset R-GCNs+XGBoost model and the comparison data.

9. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the regional economic spatiotemporal knowledge graph indicator prediction method as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the indicator prediction method for the regional economic spatiotemporal knowledge graph as described in any one of claims 1-7.

Citation Information

Cited By

  • Diffusion model training method and device based on spatial knowledge graph guidance, spatio-temporal data generation method and device, equipment and medium

    CN121436072A