Method and apparatus for predicting global telecommunication traffic demand
By constructing a global prediction model based on a geographic weighted regression tree model and integrating multi-source ground data, the problems of cross-regional migration and high-resolution modeling of communication demand models in existing technologies have been solved, enabling accurate prediction and dynamic resource scheduling of global telecommunications business communication demand.
Patent Information
- Application Number
- CN202511499224.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-20
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2045-10-20
AI Technical Summary
Existing communication demand models lack data-driven self-learning capabilities, cannot adapt to scenarios that change in different regions or times, are difficult to migrate across regions, and lack multi-source heterogeneous data fusion mechanisms, as well as high-resolution demand modeling capabilities, and cannot meet the needs of dynamic resource scheduling and interference collaborative management.
By acquiring patchwork raster data from various local regions around the world, training a geographic weighted regression tree model, and combining it with a spatial multi-fold cross-validation mechanism, a global prediction model is constructed. This model integrates multi-source ground data, including key influencing factors such as administrative divisions, population distribution, GDP, and nighttime light intensity, to achieve accurate prediction of telecommunications business communication needs.
It enables accurate prediction of global communication demand distribution covering three types of telecommunications services: terrestrial, aviation, and maritime. It improves network resource allocation efficiency and interference management capabilities, adapts to changes in different regions and times, and has high-resolution modeling capabilities.
Smart Images

Figure CN121333960B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of satellite networks and machine learning technology, and in particular to a method and apparatus for predicting communication demand in global telecommunications networks. Background Technology
[0002] Compared to traditional terrestrial networks, satellite internet services are distributed globally, and communication demands exhibit high spatial heterogeneity. This spatial heterogeneity not only affects the balanced allocation of spectrum resources and satellite link scheduling strategies but also significantly increases the complexity of cross-regional interference management. Therefore, constructing a spatial communication demand distribution model is of great significance for improving network resource allocation efficiency and supporting interference detection and cross-domain control.
[0003] To meet the demands of global network coverage and high-speed communication, the Chinese Academy of Sciences analyzed the impact of beam coverage characteristics on the performance of low-Earth orbit (LEO) satellite systems by combining user service demand models and LEO satellite beam coverage models. Nanjing University of Posts and Telecommunications proposed a statistical service modeling and real-time prediction method for satellite IoT. This method combines the traditional 3GPP (3rd Generation Partnership Project) service model with a Markov Poisson process, considering actual influencing factors to achieve source service modeling for satellite IoT, and uses a CNNBiLSTMAdaboost network for service prediction. However, existing communication demand models often rely on simplified assumptions of random spatial distribution or single data sources, failing to effectively integrate key elements such as population distribution and economic activity intensity in geographic space. This results in an inability to accurately reveal the correlation mechanism between these elements and communication demand, making it difficult to truly reflect the demand generation patterns and spatial heterogeneity in real-world scenarios. Furthermore, the model results are highly dependent on preset parameter assumptions, lack data-driven self-learning capabilities, cannot adapt to scenarios with different regional or temporal changes, and have not yet formed a universal training strategy that can be transferred across regions, making it difficult to support large-scale system-level evaluation and optimization. In addition, existing models generally lack the ability to model high-resolution requirements at a global scale, and are not suitable for sophisticated tasks such as dynamic resource scheduling and interference coordination management.
[0004] Existing multi-source heterogeneous data fusion mechanisms also have significant shortcomings in adaptability to communication services. They lack both a dedicated unified modeling framework and a fusion strategy adapted to the characteristics of the services. At the same time, existing machine learning models often exhibit black-box characteristics, with weak interpretability of their internal feature mechanisms, making it difficult to clearly reveal the correlation between the features of multi-source heterogeneous data and communication services.
[0005] With the rapid development of low-Earth orbit satellite communication networks, accurately depicting the spatial distribution of communication demand is crucial for on-board resource scheduling and interference management. Especially with the backdrop of mega-constellations serving the globe and diversifying their services to include terrestrial broadband, aviation communications, and maritime communications, traditional modeling methods, while making some progress in terrestrial traffic modeling and regional user distribution prediction, still lack a unified modeling architecture that combines geographic information systems and deep learning. Consequently, a predictive model covering the entire globe, multiple services, and a unified scale has not yet been developed, making it difficult to meet the demands for high-precision and scalable modeling. Summary of the Invention
[0006] In view of this, embodiments of the present invention provide a method and apparatus for predicting communication demand in global telecommunications networks, in order to eliminate or improve one or more defects existing in the prior art.
[0007] One aspect of the present invention provides a method for forecasting global telecommunications service communication demand, the method comprising the following steps: Acquire first-stitched raster data, global aviation data, and global maritime data for various local regions globally. The first-stitched raster data is obtained in advance by stitching multiple raster data with regional identifiers assigned to the local regions. The multiple raster data are obtained by aligning the coordinate system and data type of multi-source ground data of the local regions in a rasterized mapping manner. The multi-source ground data includes administrative division data and key influence factor data, and each raster data includes key influence factor data of at least one administrative region in the administrative division data. The first stitched raster data of each local area is input into the global prediction model, and the global prediction model outputs the telecommunications service communication demand prediction results of each local area to obtain the global terrestrial telecommunications service communication demand prediction results. The global prediction model includes a geographic weighted regression tree model. The global aviation telecommunications service communication demand forecast results are obtained based on the global aviation data, and the global maritime telecommunications service communication demand forecast results are obtained based on the global maritime data. The global telecommunications service communication demand forecast results are obtained based on the global terrestrial telecommunications service communication demand forecast results, the global aviation telecommunications service communication demand forecast results, and the global maritime telecommunications service communication demand forecast results.
[0008] In some embodiments of the present invention, the global prediction model is obtained in advance through the following steps: A global prediction model is obtained by training a local prediction model based on a cross-regional joint training set. The cross-regional joint training set includes second-stitched raster data from various local regions globally. This second-stitched raster data is obtained by stitching multiple raster data points together with regional identifiers assigned to each local region. The multiple raster data points are obtained by aligning the coordinate systems and data types of multi-source ground data in the local region using a multi-scale rasterization mapping method. The multi-source ground data includes administrative division data, key influencing factor data, and spatially autocorrelated ground telecommunications service data. The key influencing factors are those with a maximum information coefficient greater than a preset threshold and exhibiting spatial synergistic effects. Each raster data point includes key influencing factor data for at least one administrative region from the administrative division data and spatially autocorrelated ground telecommunications service communication demand data, calculated based on the ground telecommunications service data. The local prediction model includes a geographically weighted regression tree model.
[0009] In some embodiments of the present invention, the local prediction model is obtained in advance through the following steps: A spatial multi-fold cross-validation mechanism is employed to train a pre-defined geographic weighted regression tree model based on a local region training set, resulting in a local prediction model. The local region training set comprises multiple raster data points from any local region globally. These raster data points are obtained by performing coordinate system alignment, data type alignment, feature filtering, and spatial autocorrelation analysis on multi-source ground data in the local region using a multi-scale rasterization mapping method. The multi-source ground data includes administrative division data, influencing factor data, and terrestrial telecommunications service data. Each raster data point includes key influencing factor data for at least one administrative region from the administrative division data and spatially autocorrelated terrestrial telecommunications service communication demand. The terrestrial telecommunications service communication demand is calculated based on the terrestrial telecommunications service data. The geographic weighted regression tree model is constructed based on a spatial weight matrix and the key influencing factor data in the local region. Each spatial weight in the spatial weight matrix represents the degree of geographic influence between two adjacent raster data points.
[0010] In some embodiments of the present invention, a spatial multi-fold cross-validation mechanism is employed to train a preset geographically weighted regression tree model based on a local region training set, thereby obtaining a local prediction model, including: By traversing multiple hyperparameter combinations and training a preset geographic weighted regression tree model multiple times based on multiple sub-region data samples, the average prediction error of each hyperparameter combination is obtained. The hyperparameter combination with the smallest average prediction error is taken as the optimal hyperparameter combination. The multiple sub-region data samples are obtained by dividing the region formed by multiple raster data at various scales in the local area training set into multiple continuous and non-overlapping sub-regions. The hyperparameter combination includes Gaussian kernel bandwidth parameter, maximum tree depth and minimum number of leaf node samples. The average prediction error is obtained by averaging multiple prediction errors obtained during multiple training rounds of the corresponding hyperparameter combination. The local region training set and the optimal hyperparameter combination are used to train a preset geographically weighted regression tree model to obtain a local prediction model; wherein, in each round of training, Among the multiple sub-region data samples, one sub-region data sample is used as the validation set, and the other sub-region data samples are used as the training dataset. The sub-region data samples used as the validation set are different in each round. By traversing various hyperparameter combinations, a preset geographic weighted regression tree model is trained based on the training dataset. The trained geographic weighted regression tree model is then validated based on the validation set, and the prediction errors corresponding to each hyperparameter combination obtained on the validation set are recorded.
[0011] In some embodiments of the present invention, the terrestrial telecommunications service data includes the population of each administrative region, Internet access traffic, number of Internet of Things (IoT) users, and average IoT traffic per person.
[0012] In some embodiments of the present invention, the key influencing factor data includes population distribution data, GDP data, and nighttime light intensity data.
[0013] In some embodiments of the present invention, the influence factor data in the raster data is a uniform dimension influence factor feature vector.
[0014] Another aspect of the present invention provides a global telecommunications service communication demand forecasting apparatus, the apparatus comprising: a computer device including a processor and a memory, the memory storing computer instructions, the processor executing the computer instructions stored in the memory, and the apparatus implementing the steps of the aforementioned method when the computer instructions are executed by the processor.
[0015] Another aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the aforementioned method.
[0016] Another aspect of the present invention provides a computer program product including computer instructions that, when executed by a processor, implement the steps of the aforementioned method.
[0017] The global telecommunications service communication demand forecasting method and apparatus of the present invention can achieve accurate forecasting of the global communication demand distribution covering three types of telecommunications service scenarios: terrestrial, aviation, and maritime.
[0018] Additional advantages, objects, and features of the invention will be set forth in part in the description which follows, and will also become apparent in part to those skilled in the art upon studying the description, or may be learned by practice of the invention. The objects and other advantages of the invention can be realized and obtained by means of the structures specifically pointed out in the description and drawings.
[0019] Those skilled in the art will understand that the objectives and advantages achievable with the present invention are not limited to those specifically described above, and that the above and other objectives achievable with the present invention will become clearer from the following detailed description. Attached Figure Description
[0020] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, are not intended to limit the scope of the invention. The components in the drawings are not drawn to scale but are merely illustrative of the principles of the invention. For ease of illustration and description of certain parts of the invention, corresponding portions in the drawings may be enlarged, i.e., may appear larger relative to other components in an exemplary device actually manufactured according to the invention. In the drawings: Figure 1 This is a flowchart illustrating a global telecommunications service communication demand forecasting method according to an embodiment of the present invention. Figure 2 This is a schematic diagram illustrating the specific process of a global telecommunications service communication demand forecasting method in one embodiment of the present invention; Figure 3 This is a schematic diagram of the data type alignment process in one embodiment of the present invention. Detailed Implementation
[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the embodiments and accompanying drawings. Here, the illustrative embodiments and descriptions of this invention are used to explain the invention, but are not intended to limit the invention.
[0022] It should also be noted that, in order to avoid obscuring the invention with unnecessary details, only the structures and / or processing steps closely related to the solution according to the invention are shown in the accompanying drawings, while other details that are not closely related to the invention are omitted.
[0023] It should be emphasized that the term "including / comprises" as used herein refers to the presence of a feature, element, step, or component, but does not exclude the presence or addition of one or more other features, elements, steps, or components.
[0024] It should also be noted that, unless otherwise specified, the term "connection" in this article can refer not only to a direct connection, but also to an indirect connection involving an intermediary.
[0025] In the following description, embodiments of the invention will be illustrated with reference to the accompanying drawings. In the drawings, the same reference numerals represent the same or similar parts, or the same or similar steps.
[0026] This invention proposes a method and apparatus for predicting communication demand in global telecommunications networks, providing a systematic, generalizable, and multi-service modeling-based spatial prediction scheme for communication demand in telecommunications networks. The key technical problems to be solved are as follows: (1) Modeling Architecture Design for Multiple Services: In the context of global service and multi-service collaborative development of low-Earth orbit constellation systems, communication demands exhibit significant differences and complexities across three scenarios: terrestrial broadband, aviation, and maritime communications. Different service types have their own independent spatial distribution characteristics and driving mechanisms. Terrestrial services rely more on population and economic factors, aviation services are closely related to route structure and high-altitude traffic density, and maritime services are affected by port distribution and shipping routes. Therefore, designing a modeling architecture for different services is a crucial step in constructing a spatial distribution model of communication demands.
[0027] (2) Multi-source data fusion mechanism and feature interpretability modeling: The spatial distribution of communication demand is driven by a variety of heterogeneous influencing factors, such as population density, nighttime light intensity, and GDP. Different data sources have significant differences in format, resolution, coordinate system, etc., which limits the input basis of the multi-source data fusion model. At the same time, the spatial distribution of communication demand is highly nonlinear, and there are complex interactions between the features of influencing factors. It is difficult to characterize the driving relationship between variables by linear analysis alone. Although deep learning models have superior predictive performance, they are still "black boxes" in terms of feature contribution and regional interpretation. It is necessary to introduce mechanisms such as variable spatial correlation to quantitatively reveal the relationship between key influencing factors and communication demand, and to construct a high-performance and interpretable fusion modeling scheme.
[0028] like Figure 1 and Figure 2 As shown, this global telecommunications service communication demand forecasting method includes the following steps: Step S110: Obtain first stitched raster data, global aviation data, and global maritime data for various local regions globally. The first stitched raster data is obtained in advance by stitching multiple raster data with regional identifiers assigned to the local regions. The multiple raster data are obtained by aligning the coordinate system and data type of the multi-source ground data of the local regions in a rasterized mapping manner. The multi-source ground data includes administrative division data and key influence factor data, and each raster data includes key influence factor data of at least one administrative region in the administrative division data.
[0029] Specifically, "various regions around the world" can refer to different geographical areas, such as various countries, worldwide. Aviation business involves route structure, airport distribution, and flight information. Specifically, global aviation data includes raster data such as the number of routes and daily flights in various regions worldwide, aircraft types, and passenger capacity requirements for different aircraft types. Maritime business involves shipping routes, port distribution, and vessel information. Specifically, global maritime data includes raster data such as the number of routes and daily flights in various regions worldwide, vessel types, and passenger capacity requirements for different vessel types.
[0030] Multi-source ground data for various localities worldwide are obtained from multiple relevant databases in these regions, which are open to the public globally. These databases use different data formats, resulting in heterogeneous, multi-source data. For example, if a locality is within China, its administrative division data may include provincial and municipal administrative division data. These provincial and municipal administrative division data respectively include data on each provincial and municipal administrative region within China. The data source is the National Geomatics Center of China, the data type is shapefile (vector data, presented similarly to a domestic map, with attribute tables showing the province or city of each area), and the coordinate system is GCS_WGS_1984 (1984 World Geodetic System).
[0031] Key influencing factor data was determined in advance through feature screening of various influencing factors. Specifically, it was selected from numerous influencing factor data affecting the communication demand of terrestrial telecommunications services, choosing those with a maximum information coefficient greater than a preset threshold and exhibiting spatial synergistic effects—in other words, effective influencing factor data with a significant impact on the communication demand of terrestrial telecommunications services. Key influencing factor data includes population distribution data, GDP data, and nighttime light intensity data. Population distribution data can be raster data (TIF format) obtained from the global population distribution database LandScan, with the data coordinate system being the GCS_WGS_1984 geographic coordinate system. GDP data can be TIF raster data obtained from publicly available journal articles, with the data coordinate system being the WGS 1984 Albers Conical Equal Area (WGS_1984_Albers) coordinate system. Nighttime light intensity data can be corrected nighttime light intensity TIF raster data obtained from publicly available websites, with the data coordinate system being the WGS 1984 Albers projection (WGS_1984_Albers) coordinate system. The aforementioned TIFF raster data is presented as multiple raster blocks, and its attribute table displays the raster number and the population, GDP, or nighttime light intensity values stored within each raster. It should also be noted that all the multi-source ground data mentioned above are from the same valid time period.
[0032] Because the acquired multi-source ground data are located in different coordinate systems, to ensure consistency in subsequent fusion modeling steps, it is necessary to spatially align and map the heterogeneous data on administrative divisions, population distribution, GDP, and nighttime light intensity to a unified two-dimensional spatial system. This means that for various local regions globally, the aforementioned multi-source ground data undergoes coordinate system and data type alignment preprocessing using raster mapping to obtain corresponding raster data. Specifically, this includes: firstly, mapping all multi-source ground data to a two-dimensional plane using the Mercator projection on the GIS platform to obtain the spatial extent of the coordinate-aligned multi-source ground data, denoted as . This allows different data to be overlaid and statistically analyzed under the same spatial reference. Based on the set spatial scale parameters... ( (This refers to a specific scale, such as 10km or 100km resolution, etc.) to define the spatial range. Divided into The raster array is used to construct a unified global raster index, in which... Indicates the total number of grid cells in the row direction. This represents the total number of grid cells along the column direction, with each grid cell denoted as . , , and These represent the raster numbers in the row and column directions, respectively, and each raster cell is associated with its corresponding actual latitude and longitude. The latitude and longitude are the center points of the corresponding raster cells. This transforms multi-source ground data into a raster dataset with the same spatial area, aligned coordinate system, and rasterized structure. The attribute table of this dataset includes the raster number and the latitude and longitude of each raster cell. Next, all feature data after coordinate system alignment are aligned by data type to establish corresponding raster layers. Specifically, this is done by performing partitioned statistics on each raster cell, mapping administrative division data, population distribution data, GDP data, and nighttime light intensity data to the corresponding raster cells, forming corresponding administrative division layers, population raster layers, GDP raster layers, and nighttime light intensity layers. Correspondingly, the attribute table also includes the administrative region (province or city), population, GDP value, and nighttime light intensity value of each raster cell. Thus, the administrative division data and various key influencing factor data are aligned with the latitude and longitude of each raster cell, resulting in key influencing factor data. , can be represented as: in, , and These can be represented by population, GDP, and nighttime light intensity, respectively, in the [missing data]. The values within each raster cell. That is, the population, GDP, and nighttime light intensity of each province or city in each raster cell are known. Thus, global ground raster data is obtained.
[0033] Assign corresponding unique numbers to each local area. Each region identifier serves as a region identifier, and these region identifiers are assigned to all raster cells within each local area. Specifically, for each region identifier, one-hot encoding is first performed to obtain the region embedding vector. Thus, the region embedding matrix formed by the region embedding vectors is obtained. The region embedding vector and key influence factor data are concatenated according to the following formula to obtain the first concatenated raster data. : in, This indicates a splicing operation. Indicates the encoding length. Indicates the number of local regions. Indicates the serial number of a local area.
[0034] In some embodiments, before concatenation, the method further standardizes the key influencing factor data to eliminate the influence of different factor data units. Specifically, Z-score standardization can be performed on all raster cell data in the three categories of population, GDP, and nighttime light intensity according to the following formula to obtain the feature vector of the key influencing factor, which is then concatenated with the corresponding region embedding vector: in, and They are The mean and standard deviation, Indicates the sequence number of the key influencing factor.
[0035] Step S120: Input the first stitched raster data of each local area into the global prediction model, and output the telecommunications service communication demand prediction results of each local area by the global prediction model to obtain the global terrestrial telecommunications service communication demand prediction results. The global prediction model includes a geographic weighted regression tree model.
[0036] In some embodiments, the global prediction model is trained in advance through the following steps: A global prediction model is obtained by training a local prediction model based on a cross-regional joint training set. The cross-regional joint training set includes second-stitched raster data from various local regions globally. This second-stitched raster data is obtained by stitching multiple raster data points together with regional identifiers assigned to each local region. The multiple raster data points are obtained by aligning the coordinate systems and data types of multi-source ground data in the local region using a multi-scale rasterization mapping method. The multi-source ground data includes administrative division data, key influencing factor data, and spatially autocorrelated ground telecommunications service data. The key influencing factors are those with a maximum information coefficient greater than a preset threshold and exhibiting spatial synergistic effects. Each raster data point includes key influencing factor data for at least one administrative region from the administrative division data and spatially autocorrelated ground telecommunications service communication demand data, calculated based on the ground telecommunications service data. The local prediction model includes a geographically weighted regression tree model.
[0037] In some embodiments, the local prediction model is trained in advance through the following steps: A spatial multi-fold cross-validation mechanism is employed to train a pre-defined geographic weighted regression tree model based on a local region training set, resulting in a local prediction model. The local region training set comprises multiple raster data points from any local region globally. These raster data points are obtained by performing coordinate system alignment, data type alignment, feature filtering, and spatial autocorrelation analysis on multi-source ground data in the local region using a multi-scale rasterization mapping method. The multi-source ground data includes administrative division data, influencing factor data, and terrestrial telecommunications service data. Each raster data point includes key influencing factor data for at least one administrative region from the administrative division data and spatially autocorrelated terrestrial telecommunications service communication demand. The terrestrial telecommunications service communication demand is calculated based on the terrestrial telecommunications service data. The geographic weighted regression tree model is constructed based on a spatial weight matrix and the key influencing factor data in the local region. Each spatial weight in the spatial weight matrix represents the degree of geographic influence between two adjacent raster data points.
[0038] Specifically, taking any local region globally as an example, the implementation process of the above steps will be explained. Other regions globally can also be selected for research. The acquisition method for multi-source terrestrial data in the domestic region is similar to the acquisition process described in the previous embodiments, and will not be repeated here. Terrestrial telecommunications service data for the domestic region can be obtained from statistical yearbook PDF documents published by the National Bureau of Statistics. It is important to note that the terrestrial telecommunications service data should be within the same valid time period as the acquired administrative division data and influencing factor data. This facilitates the accurate identification of the correlation between various influencing factors and terrestrial telecommunications service communication needs.
[0039] The influencing factor data in this step includes various types of data that may affect the communication demand of terrestrial telecommunications services, such as population distribution data, GDP data, nighttime light intensity data, and ground slope. Terrestrial telecommunications service data includes population size, internet access traffic, number of IoT users, and average IoT data traffic per person in each administrative region.
[0040] Next, multi-scale rasterization mapping preprocessing is performed on the multi-source ground data of the domestic region. This process is similar to the rasterization mapping process described in the previous embodiment. One difference is that in the coordinate system alignment process of this embodiment, it is based on multiple set spatial scale parameters. Multiple scales of rasters are obtained to obtain raster data aligned with coordinate systems of multiple scales, thereby increasing the data diversity of the training set, whereas the aforementioned embodiment uses a single scale for raster division.
[0041] The second difference lies in the data type alignment in this embodiment, as follows: Figure 3The process is as follows. Specifically, for each raster scale, the following processing is performed: First, administrative division data, population distribution data, GDP data, nighttime light intensity data, and ground slope data, as well as other influencing factors, are mapped to the corresponding raster cells, forming corresponding layers (the attribute table also contains the ground slope and other influencing factor values and administrative region data for each raster cell), thus obtaining the influencing factor data. , can be represented as: in, Indicates the number of impact factors. Indicates the first The impact factor in the first The values in each raster cell. That is, to obtain data such as population, GDP, nighttime light intensity, and ground slope for each province or city in each raster cell at each scale.
[0042] Secondly, a correlation model between terrestrial telecommunications service communication demand and population density is constructed. Specifically, terrestrial telecommunications service data and administrative division layers are first linked to generate spatial provincial indicators (terrestrial telecommunications service data is upgraded to a terrestrial telecommunications service raster layer, i.e., obtaining data such as population, internet access traffic, number of IoT users, and average IoT traffic per capita for each province in each raster cell). The terrestrial telecommunications service raster layer and the population raster layer are then linked (terrestrial telecommunications service data is updated in the attribute table of the raster data), enabling each raster cell to display various administrative regions and their populations, internet access traffic, number of IoT users, and average IoT traffic per capita, among other service data. If multiple administrative regions exist in a raster cell, the relevant data for each administrative region are listed in the raster cell. Then, the terrestrial telecommunications service communication demand indicator data for each administrative region in each raster cell is calculated to obtain the terrestrial telecommunications service communication demand for each raster cell. The terrestrial telecommunications service communication demand indicator data includes internet traffic and IoT traffic, and the calculation formulas are as follows: For each grid cell, if there are multiple administrative regions within the grid cell, the terrestrial telecommunications service communication demand (the sum of Internet service volume and Internet of Things service volume) of these multiple administrative regions is combined and used as the terrestrial telecommunications service communication demand of that grid cell. This ultimately yields a raster layer depicting the communication demand for terrestrial telecommunications services. Thus, administrative division data, various influencing factor data, and terrestrial telecommunications service communication demand data are aligned within a unified two-dimensional raster plane. In other words, for raster data at various scales, the values of various influencing factors and the communication demand for terrestrial telecommunications services for each administrative region within each raster cell are obtained.
[0043] Optionally, to eliminate the influence of different data units, this method also processes the aforementioned influencing factor data. and terrestrial telecommunications communication demand Standardization is performed to obtain corresponding feature vectors for each factor, where the feature vector of the influence factor can be expressed as follows: The processing procedure is similar to that in the aforementioned embodiments and will not be repeated here. Therefore, in the following formulas, correspondingly, All can be replaced with , All can be replaced with , Alternatively, all of them can be replaced with feature vectors of communication demand for terrestrial telecommunications services. It is the average value obtained from the feature vector of communication demand for terrestrial telecommunications services.
[0044] The following section describes feature selection for various influencing factor data and spatial autocorrelation analysis of terrestrial telecommunications service communication demand. Feature selection includes maximum information coefficient (MIC) analysis and spatial synergy analysis. Specifically, the MIC is used to analyze and identify influencing factors with potentially nonlinear or complex functional relationships based on various influencing factor data. and Mapped onto a two-dimensional raster plane, for various influencing factors, the mutual information value between the influencing factors and the communication demand of terrestrial telecommunications services under all possible (multiple scales) raster divisions is calculated, and the MIC value is obtained by normalizing the information value. The calculation formula can be expressed as: in, Represents grid division The mutual information value between the lower influence factor and the communication demand of terrestrial telecommunications services. This represents the set of all possible raster cells for partitioning (including multi-scale raster data). If , If the set MIC threshold is not met, the influencing factor is retained and a candidate key feature set is formed. A value of 0.25 is acceptable. In other words, through maximum information coefficient analysis, from... Retaining from the impact factors Influencing factors .
[0045] Moran's I index is used to analyze the spatial autocorrelation of communication demand in terrestrial telecommunications services and verify its spatial structure characteristics. The formula can be expressed as: in, Represents grid cells and its adjacent grid cells The spatial weight matrix elements between them This represents the sum of spatial weights. This represents the average communication demand for terrestrial telecommunications services across all grid cells. (The quantified value of terrestrial telecommunications service communication demand is...) If the demand for terrestrial telecommunications services exceeds a preset threshold, it demonstrates positive spatial autocorrelation. The results show that the demand for terrestrial telecommunications services exhibits significant positive spatial autocorrelation within the study area, demonstrating significant spatial clustering. This makes it suitable for modeling using a model that captures spatial heterogeneity, such as a geographically weighted regression tree model. If the demand for terrestrial telecommunications services lacks spatial autocorrelation, other applicable models are used.
[0046] Furthermore, bivariate spatial autocorrelation analysis is introduced to identify influencing factors with spatial synergistic effects based on candidate key feature sets: in, This represents the average value of the influencing factor data across all raster cells for each influencing factor selected as a key feature in the candidate key feature set. Thus, through feature selection, population distribution, GDP, nighttime light intensity, and ground slope are identified as key influencing factors. Consequently, a local region training set is obtained.
[0047] Based on the above scheme, the linear or nonlinear strength of the relationship between various influencing factor characteristics and the target communication demand intensity is assessed using MIC (Microscopic Interference Model). The spatial synergistic effect between these two relationships is assessed using Moran's I and bivariate spatial autocorrelation analysis. Based on these three statistical methods, the correlation between various influencing factor characteristics and communication demand is quantitatively analyzed.
[0048] In one embodiment, a geographic weighted regression tree model is used for modeling. A spatial weight matrix is constructed: spatial neighborhood influence weights are built using a Gaussian kernel function to capture the spatial heterogeneity of geographic influence and communication needs between adjacent raster cells. The spatial weights of any two raster cell data can be defined as follows: in, Represents grid with his neighbors The center distance, This represents the Gaussian kernel bandwidth parameter, used to control the rate of spatial weight decay. The bandwidth setting directly affects the model's sensitivity to local spatial features. The smaller the bandwidth, the faster the weight decays, and the more sensitive the model becomes. The optimal value needs to be determined through spatial cross-validation.
[0049] Model Structure Design: GWRT (Geographically Weighted Regression Tree) is an ensemble learning model that embeds the concept of geographic weighting into a tree structure. While retaining the flexible nonlinear expression capabilities of traditional decision trees, it also possesses stronger regional adaptability, enabling it to capture the changing relationships between different key influencing factors and communication needs within a local area. The constructed geographically weighted regression tree prediction model can be represented as follows: in, This represents the predicted demand for terrestrial telecommunications services. Represents the spatial weight matrix. This represents a regression tree that incorporates a spatial weighting mechanism.
[0050] The splitting criterion for each node in the tree structure is driven by the weighted mean squared error (MSE): in, This represents a local training set, which includes raster data at multiple scales. Each scale raster data includes multiple raster cell data, and each raster cell data includes population, GDP value, nighttime light intensity value, and actual value of terrestrial telecommunications service communication demand. Represents grid The neighborhood space or set of neighbors, Representing neighbor grids The true value of terrestrial telecommunications service communication demand. Representing neighbor grids Forecast values of terrestrial telecommunications service communication demand.
[0051] Specifically, a spatial multi-fold cross-validation mechanism is used to train a pre-defined geographically weighted regression tree model based on a local region training set to obtain a local prediction model, including the following steps: By traversing multiple hyperparameter combinations and training a preset geographic weighted regression tree model multiple times based on multiple sub-region data samples, the average prediction error of each hyperparameter combination is obtained. The hyperparameter combination with the smallest average prediction error is taken as the optimal hyperparameter combination. The multiple sub-region data samples are obtained by dividing the region formed by multiple raster data at various scales in the local area training set into multiple continuous and non-overlapping sub-regions. The hyperparameter combination includes Gaussian kernel bandwidth parameter, maximum tree depth and minimum number of leaf node samples. The average prediction error is obtained by averaging multiple prediction errors obtained during multiple training rounds of the corresponding hyperparameter combination. The local region training set and the optimal hyperparameter combination are used to train a preset geographically weighted regression tree model to obtain a local prediction model; wherein, in each round of training, Among the multiple sub-region data samples, one sub-region data sample is used as the validation set, and the other sub-region data samples are used as the training dataset. The sub-region data samples used as the validation set are different in each round. By traversing various hyperparameter combinations, a preset geographic weighted regression tree model is trained based on the training dataset. The trained geographic weighted regression tree model is then validated based on the validation set, and the prediction errors corresponding to each hyperparameter combination obtained on the validation set are recorded.
[0052] Specifically, the spatial multi-fold cross-validation mechanism is a spatial K-fold cross-validation mechanism. During K rounds of model training and hyperparameter optimization using this mechanism, the model is divided into K continuous and non-overlapping sub-regions. The optimal hyperparameter combination... It can be represented as: in, This indicates the maximum depth of the regression tree, used to control model complexity; This represents the minimum number of samples for each leaf node; Indicates the index of the combination of hyperparameters for traversal. This indicates the number of hyperparameter combinations being iterated over. Indicates the first The prediction error of a combination of hyperparameters.
[0053] After training, the importance contribution of each key influencing factor's features is calculated by statistically analyzing the number of splits each input variable (key influencing factor) participates in within the regression tree structure and the resulting error reduction, and by accumulating the split gain. This allows for the evaluation and interpretation of the impact of each key influencing factor on communication requirements. The calculation method for feature importance contribution can be as follows: in, Characteristics of key influencing factors The number of times a tree participates in splitting. Indicates the index of the optimal hyperparameter combination. This represents the decrease in weighted error caused by the split. The local prediction model obtained through training can fuse multi-source influencing factors in a local area to predict the communication demand of telecommunications services within that area.
[0054] To enhance the model's adaptability and generalization ability to different geographical regions, multiple representative local areas were selected from around the world. Multi-source ground data from these local areas were used to construct a cross-regional joint training set, which was then used to train the local prediction model across multiple regions to obtain the global prediction model.
[0055] By identifying key influencing factors and spatial autocorrelation analysis, and determining the communication demand of terrestrial telecommunications services with spatial autocorrelation, it becomes clear which terrestrial telecommunications service data exhibit spatial autocorrelation. Therefore, when acquiring multi-source terrestrial data, the key influencing factor data and spatially autocorrelation-related terrestrial telecommunications service data are directly obtained. For multi-source terrestrial data in various local areas, preprocessing is performed similarly to the multi-scale rasterization mapping method used for domestic regional multi-source terrestrial data in the aforementioned example to obtain multi-scale raster data. The stitching method for obtaining the second stitched raster data is the same as the stitching method for obtaining the first stitched raster data described in the aforementioned embodiment, and will not be repeated here. By introducing a learnable regional embedding vector for each local area... This forms a regional embedding matrix, which serves as a new trainable parameter. This enables the trained model to perceive differences in geographical regional structures, allowing it to adaptively adjust the prediction mechanism and weight allocation within different regions, thus improving its ability to model regional heterogeneity. Therefore, by introducing a regional embedding perception mechanism, the global prediction model obtained after training can be transferred to a global scale for predicting the distribution of global telecommunications network service communication demands.
[0056] Step S130: Based on the global aviation data, estimate the global aviation telecommunications service communication demand forecast results; based on the global maritime data, estimate the global maritime telecommunications service communication demand forecast results.
[0057] Specifically, based on the number of routes and the total number of daily flights within the global aviation data statistical grid, and combined with the aircraft type and passenger capacity demand estimates for different aircraft types, a global aviation telecommunications service communication demand forecast is generated, and this forecast is displayed in an aviation business database. Similarly, based on the number of routes and the total number of daily flights within the global maritime data statistical grid, and combined with the ship type and passenger capacity demand estimates for different ship types, a global maritime telecommunications service communication demand forecast is generated, and this forecast is displayed in a maritime business database.
[0058] Step S140: Based on the global terrestrial telecommunications service communication demand forecast results, the global aviation telecommunications service communication demand forecast results, and the global maritime telecommunications service communication demand forecast results, obtain the global telecommunications service communication demand forecast results.
[0059] Specifically, by projecting the communication demand forecasts for global terrestrial, aeronautical, and maritime telecommunications networks onto a unified grid system, the three types of results can be weighted and fused according to the following expression to obtain the final global telecommunications network service communication demand forecast: in, The results of global telecommunications demand forecasts are presented in the form of a map database. This represents the forecast results for global terrestrial telecommunications demand. This indicates the forecast results for global aviation telecommunications business communication demand. This section presents the global maritime telecommunications demand forecast results, while the aviation and maritime communication demand forecast results are displayed in database format. , and Let represent the fusion weight coefficients, and satisfy: , The fusion weighting coefficient can be set as a fixed value, a regional variable parameter, or a dynamic adjustment function.
[0060] A structured, predicted global telecommunications network service communication demand map database is established using a GIS platform, and data storage and spatial visualization are performed. In some embodiments, the predicted communication demand results for three types of global telecommunications services—terrestrial, aviation, and maritime—can be visualized independently; in other embodiments, the fused result, i.e., the global telecommunications service communication demand prediction result, can also be visualized. The core fields for visualization include: grid ID and center point latitude and longitude, key influencing factor values, predicted values for terrestrial, aviation, and maritime telecommunications service communication demand, the final fused communication demand prediction result, and a description of the model version and fusion strategy.
[0061] In summary, the global telecommunications demand forecasting method proposed in this invention addresses the challenges of modeling the spatial heterogeneity and realism of communication demand in low-Earth orbit satellite communication systems. By integrating multi-source global terrestrial, aeronautical, and maritime data—including socioeconomic data such as population, GDP, and nighttime light intensity that influence communication demand—a multi-dimensional spatial dataset covering terrestrial, aeronautical, and maritime sectors is constructed. This overcomes the limitations of traditional communication modeling, which relies on random distribution assumptions or single data sources and suffers from insufficient spatial representation. Furthermore, by constructing a unified grid system and a multi-source influencing factor fusion mechanism, the above-mentioned influencing factor data are fused from multiple sources, improving the realism of terrestrial communication demand distribution modeling and accurately depicting the spatial heterogeneity of terrestrial communication demand. This enables accurate prediction of global communication demand distribution across terrestrial, aeronautical, and maritime telecommunications service scenarios.
[0062] Addressing the challenges of a lack of multi-source fusion mechanisms and insufficient interpretability of traditional machine learning models in the field of communication demand forecasting, this invention introduces a multi-source spatial data fusion mechanism into the field of LEO communication demand forecasting for the first time. It proposes a geographically weighted regression tree model to achieve interpretable modeling of communication demand. This model introduces neighborhood influence weights through a spatial Gaussian kernel function and uses a geographically weighted regression tree to model the local nonlinear relationship between multi-source influence factors and communication demand. Combined with spatial cross-validation to optimize model parameters, it can not only accurately fit the nonlinear distribution of communication demand but also output the feature contribution of multi-source influence factors, thus achieving a dual guarantee of accuracy in communication load forecasting and interpretability of influence factors. By employing a regional embedding perception mechanism to adapt to the spatial heterogeneity of different regions, the model exhibits good generalization ability and can be applied to communication demand forecasting in different geographic spaces globally. Furthermore, by designing a service-weighted fusion mechanism to generate a unified prediction map, it supports the visualization of the spatial distribution of communication demand globally.
[0063] Corresponding to the above method, the present invention also provides a global telecommunications service communication demand forecasting device, the device including a computer device, the computer device including a processor and a memory, the memory storing computer instructions, the processor being used to execute the computer instructions stored in the memory, and when the computer instructions are executed by the processor, the device implements the steps of the aforementioned method.
[0064] This invention also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the aforementioned method. The computer-readable storage medium may be a tangible storage medium, such as random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, floppy disk, hard disk, removable storage disk, CD-ROM, or any other form of storage medium known in the art.
[0065] This invention also provides a computer program product, including computer instructions that, when executed by a processor, implement the steps of the aforementioned method.
[0066] Those skilled in the art will understand that the exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Whether implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this invention. When implemented in hardware, it can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this invention are programs or code segments used to perform the desired tasks. The programs or code segments can be stored in a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried in a carrier wave.
[0067] It should be clarified that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of the present invention.
[0068] In this invention, features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, and / or combined with or in place of features of other embodiments.
[0069] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, various modifications and variations of the embodiments of the present invention are possible. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for forecasting global telecommunications service communication demand, characterized in that, The method includes: Acquire first-stitched raster data, global aviation data, and global maritime data for various local regions globally. The first-stitched raster data is obtained in advance by stitching multiple raster data with regional identifiers assigned to the local regions. The multiple raster data are obtained by aligning the coordinate system and data type of multi-source ground data of the local regions in a rasterized mapping manner. The multi-source ground data includes administrative division data and key influence factor data, and each raster data includes key influence factor data of at least one administrative region in the administrative division data. The first stitched raster data of each local region is input into the global prediction model, which outputs the telecommunications service communication demand prediction results for each local region to obtain the global terrestrial telecommunications service communication demand prediction results. The global prediction model includes a geographic weighted regression tree model. The global prediction model is pre-trained through the following steps: training the local prediction model based on a cross-regional joint training set to obtain the global prediction model. The cross-regional joint training set includes second stitched raster data from various local regions globally. The second stitched raster data is obtained by stitching multiple raster data with regional identifiers assigned to each local region. This method involves aligning the coordinate system and data types of multi-source ground data for the local area using a multi-scale rasterized mapping approach. The multi-source ground data includes administrative division data, key influencing factor data, and spatially autocorrelated ground telecommunications service data. Key influencing factors are those with a maximum information coefficient greater than a preset threshold and exhibiting spatial synergistic effects. Each raster data unit includes key influencing factor data for at least one administrative region from the administrative division data and spatially autocorrelated ground telecommunications service communication demand data, calculated based on the ground telecommunications service data. The local prediction model includes a geographically weighted regression tree model. The global aviation telecommunications service communication demand forecast results are obtained based on the global aviation data, and the global maritime telecommunications service communication demand forecast results are obtained based on the global maritime data. The global telecommunications service communication demand forecast results are obtained based on the global terrestrial telecommunications service communication demand forecast results, the global aviation telecommunications service communication demand forecast results, and the global maritime telecommunications service communication demand forecast results.
2. The method according to claim 1, characterized in that, The local prediction model is pre-trained through the following steps: A spatial multi-fold cross-validation mechanism is employed to train a pre-defined geographic weighted regression tree model based on a local region training set, resulting in a local prediction model. The local region training set comprises multiple raster data points from any local region globally. These raster data points are obtained by performing coordinate system alignment, data type alignment, feature filtering, and spatial autocorrelation analysis on multi-source ground data in the local region using a multi-scale rasterization mapping method. The multi-source ground data includes administrative division data, influencing factor data, and terrestrial telecommunications service data. Each raster data point includes key influencing factor data for at least one administrative region from the administrative division data and spatially autocorrelated terrestrial telecommunications service communication demand. The terrestrial telecommunications service communication demand is calculated based on the terrestrial telecommunications service data. The geographic weighted regression tree model is constructed based on a spatial weight matrix and the key influencing factor data in the local region. Each spatial weight in the spatial weight matrix represents the degree of geographic influence between two adjacent raster data points.
3. The method according to claim 2, characterized in that, A spatial multi-fold cross-validation mechanism is employed to train a pre-defined geographically weighted regression tree model based on a local region training set, resulting in a local prediction model, including: By traversing multiple hyperparameter combinations and training a preset geographic weighted regression tree model multiple times based on multiple sub-region data samples, the average prediction error of each hyperparameter combination is obtained. The hyperparameter combination with the smallest average prediction error is taken as the optimal hyperparameter combination. The multiple sub-region data samples are obtained by dividing the region formed by multiple raster data at various scales in the local area training set into multiple continuous and non-overlapping sub-regions. The hyperparameter combination includes Gaussian kernel bandwidth parameter, maximum tree depth and minimum number of leaf node samples. The average prediction error is obtained by averaging multiple prediction errors obtained during multiple training rounds of the corresponding hyperparameter combination. The local region training set and the optimal hyperparameter combination are used to train a preset geographically weighted regression tree model to obtain a local prediction model; wherein, in each round of training, Among the multiple sub-region data samples, one sub-region data sample is used as the validation set, and the other sub-region data samples are used as the training dataset. The sub-region data samples used as the validation set are different in each round. By traversing various hyperparameter combinations, a preset geographic weighted regression tree model is trained based on the training dataset. The trained geographic weighted regression tree model is then validated based on the validation set, and the prediction errors corresponding to each hyperparameter combination obtained on the validation set are recorded.
4. The method according to claim 1 or 2, characterized in that, The terrestrial telecommunications service data includes the population of each administrative region, internet access traffic, number of IoT users, and average IoT traffic per person.
5. The method according to any one of claims 1 to 2, characterized in that, The key influencing factor data includes population distribution data, GDP data, and nighttime light intensity data.
6. The method according to claim 1, characterized in that, The impact factor data in the raster data is a uniform dimension impact factor feature vector.
7. A global telecommunications service communication demand forecasting device, comprising a processor, a memory, and computer instructions stored in the memory, characterized in that, The processor is configured to execute the computer instructions, and when the computer instructions are executed, the device implements the steps of the method as described in any one of claims 1 to 6.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method as described in any one of claims 1 to 6.
9. A computer program product comprising computer instructions, characterized in that, When executed by a processor, the computer instructions implement the steps of the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-element composite weighted estimation method for broadband low-orbit satellite communication traffic
CN112751604A