Traffic prediction method and device in wireless communication, electronic equipment and storage medium

By extracting multi-scale temporal features and spatial dependency modeling, generating node embedding vectors, and selecting target expert models for prediction, the problems of multi-scale information being difficult to learn and heterogeneous interaction relationships being difficult to model in existing technologies are solved, thereby improving the prediction accuracy of spatiotemporal communication traffic.

CN120640333AActive Publication Date: 2025-09-12SHENZHEN RES INST OF BIG DATA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511087894.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-09-12
Estimated Expiration
2045-08-05

AI Technical Summary

Technical Problem

Existing technologies have difficulty in effectively learning multi-scale information and modeling the complex heterogeneous interaction relationships between multi-scale temporal patterns in different spatial nodes when processing long-term spatiotemporal data, resulting in insufficient accuracy in spatiotemporal communication traffic prediction.

Method used

By extracting multi-scale temporal features, spatial dependency modeling is performed, node embedding vectors are generated, and the target expert model is selected for prediction based on the node embedding vectors. A hybrid expert model architecture is used for feature processing.

Benefits of technology

It improves the prediction accuracy of spatiotemporal communication traffic, can effectively learn multi-scale temporal features in long-term sequences and accurately model the complex heterogeneous interaction relationships between different spatial nodes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120640333A_ABST
    Figure CN120640333A_ABST
Patent Text Reader

Abstract

The invention provides a traffic prediction method and device in wireless communication, electronic equipment and a storage medium. The method comprises the following steps: acquiring historical space-time traffic data of a plurality of space nodes in a target area; based on the historical space-time traffic data of each space node, multi-scale time features of the space nodes are extracted; spatial dependency modeling is carried out on the multi-scale time features, and a node embedding vector corresponding to each spatial node is obtained; based on the node embedding vector corresponding to each space node, determining a target expert model from a plurality of preset to-be-selected expert models for the corresponding space node; the multi-scale time features are predicted through the target expert model corresponding to each space node, target communication flow predicted values of the multiple space nodes are obtained, the multi-scale time features in a long-term sequence can be effectively learned, the complex heterogeneity interaction relation of the multi-scale time features among the different space nodes is accurately modeled, and the communication flow prediction efficiency is improved. Therefore, the prediction precision of the space-time communication flow is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of wireless communications, and in particular to a method, device, electronic device, and storage medium for traffic prediction in wireless communications. Background Art

[0002] Spatiotemporal forecasting of communication traffic is a key task in intelligent network management and resource scheduling. Its goal is to accurately predict future changes in communication traffic at different geographic locations based on historical data. With the development of 5G and future networks, network traffic exhibits significant periodicity and burstiness in the temporal dimension, and strong regional coupling and topological dependence in the spatial dimension. Therefore, effectively modeling the temporal dependencies and spatial correlations in communication traffic to achieve high-precision spatiotemporal traffic forecasting has become a critical issue for improving network efficiency and user service quality.

[0003] Graph neural networks are widely used in spatial modeling, such as spatiotemporal graph convolutional networks (GCNNs) that leverage predefined graphs or adaptive GCNs that can adaptively learn graph structures. For temporal modeling, the Mamba architecture, based on state-space models, has attracted considerable attention due to its high efficiency. However, these existing technologies suffer from two key drawbacks when processing long-term spatiotemporal data: first, the multi-scale information inherent in long-term sequences is difficult to learn simultaneously and effectively; second, the complex and heterogeneous interactions between these multi-scale temporal patterns across different spatial nodes cannot be fully modeled. Summary of the Invention

[0004] The embodiments of the present application provide a traffic prediction method, device, electronic device and storage medium in wireless communications, which can effectively learn multi-scale time features in long-term sequences and accurately model the complex heterogeneous interaction relationships between multi-scale time features between different spatial nodes, thereby improving the prediction accuracy of spatiotemporal communication traffic.

[0005] To achieve the above objectives, a first aspect of an embodiment of the present application provides a method for traffic prediction in wireless communications, the method comprising: Obtain historical spatiotemporal traffic data for multiple spatial nodes in the target area; Extracting multi-scale time features of each spatial node based on the historical spatiotemporal traffic data of the spatial node; Performing spatial dependency modeling on the multi-scale temporal features to obtain a node embedding vector corresponding to each of the spatial nodes; Based on the node embedding vector corresponding to each of the spatial nodes, determining a target expert model for the corresponding spatial node from a plurality of preset candidate expert models; The multi-scale time characteristics are predicted by the target expert model corresponding to each of the spatial nodes to obtain target communication flow prediction values ​​of multiple spatial nodes.

[0006] In some embodiments, extracting the multi-scale time features of each spatial node based on the historical spatiotemporal traffic data of the spatial node includes: Projecting the historical spatiotemporal traffic data of each spatial node into a latent space feature tensor through a multi-layer perceptron; wherein the latent space feature tensor includes a preset latent space feature dimension; Performing convolution processing on each of the latent space feature tensors through a plurality of one-dimensional convolution layers with different preset convolution kernel sizes to obtain temporal feature representations at multiple time scales; For each of the spatial nodes, the time feature representations at multiple time scales are stacked to obtain the multi-scale time feature of the spatial node.

[0007] In some embodiments, obtaining historical spatiotemporal traffic data of multiple spatial nodes in the target area includes: Obtaining device measurement data reported by all terminal devices in the target area; wherein the device measurement data includes geographic location data and traffic data; Dividing the target area into a plurality of preset geographic grids, and determining each of the geographic grids as a spatial node; Based on the geographic location data and the traffic data, the historical traffic data of each of the spatial nodes in each preset time segment is determined to obtain the historical spatiotemporal traffic data of each of the spatial nodes.

[0008] In some embodiments, performing spatial dependency modeling on the multi-scale temporal features to obtain a node embedding vector corresponding to each of the spatial nodes includes: generating an adaptive graph adjacency matrix for characterizing spatial dependency relationships between the plurality of spatial nodes based on a preset initial node embedding vector for each of the spatial nodes; Based on the adaptive graph adjacency matrix, performing graph convolution processing on the multi-scale temporal features to aggregate spatial neighborhood information of the multiple spatial nodes to obtain multi-scale temporal aggregated features; The multi-scale temporal aggregation features are cross-scale fused to obtain a node embedding vector corresponding to each of the spatial nodes.

[0009] In some embodiments, cross-scale fusing the multi-scale temporal aggregation features to obtain a node embedding vector corresponding to each of the spatial nodes includes: Projecting the multi-scale temporal aggregation features into a query matrix, a key matrix, and a value matrix respectively; Performing matrix multiplication on the query matrix and the key matrix, and scaling and masking the multiplication result to obtain an attention score; wherein the masking adopts a preset lower triangular mask; Normalizing the attention score and performing matrix multiplication with the value matrix to obtain an attention output feature; Performing residual connection and layer normalization processing on the attention output feature and the multi-scale temporal aggregation feature to generate a fused spatiotemporal feature; Based on the fused spatiotemporal features, the node embedding vector corresponding to each of the spatial nodes is determined.

[0010] In some embodiments, the plurality of preset candidate expert models are set to include a shared expert model and a plurality of dedicated expert models, and determining a target expert model from the plurality of preset candidate expert models for the corresponding spatial node based on the node embedding vector corresponding to each spatial node includes: Calculating a routing score corresponding to each of the dedicated expert models based on the node embedding vector; The dedicated expert model and the shared expert model with the highest selection routing score are jointly determined as the target expert model.

[0011] In some embodiments, predicting the multi-scale time feature by the target expert model corresponding to each of the spatial nodes to obtain target communication flow prediction values ​​of the plurality of spatial nodes includes: The multi-scale time features are respectively extracted using the selected dedicated expert model and the shared expert model to obtain dedicated features and shared features; wherein the feature extraction process includes: Performing time scale expansion processing on the multi-scale time feature through multiple one-dimensional convolution layers with different convolution kernel sizes to obtain an expanded multi-scale time feature; Performing multi-scale feature extraction processing on the expanded multi-scale temporal features through a preset learnable multi-scale bias; Fusing the dedicated features with the shared features to obtain target fused features; The target fusion features are mapped into a prediction space by a multi-layer perceptron to obtain the target communication flow prediction value.

[0012] In some embodiments, after obtaining the target communication traffic prediction values ​​of the plurality of spatial nodes, the method further includes: Based on a preset joint loss function, calculating the loss value between the target communication traffic prediction value and the corresponding actual traffic value; wherein the joint loss function includes a mean absolute error loss and a causal contrast loss, the mean absolute error loss is used to measure the prediction error, and the causal contrast loss is used to promote mode decoupling between the multiple candidate expert models; Based on the loss value, the multiple candidate expert models are updated.

[0013] To achieve the above-mentioned objectives, a second aspect of an embodiment of the present application provides a traffic prediction device in wireless communication, comprising: An acquisition module is used to obtain historical spatiotemporal traffic data of multiple spatial nodes in the target area; an extraction module, configured to extract multi-scale time features of each spatial node based on the historical spatiotemporal traffic data of the spatial node; A modeling module, configured to perform spatial dependency modeling on the multi-scale temporal features to obtain a node embedding vector corresponding to each of the spatial nodes; a determination module, configured to determine a target expert model from a plurality of preset candidate expert models for the corresponding spatial node based on the node embedding vector corresponding to each spatial node; The prediction module is used to predict the multi-scale time characteristics through the target expert model corresponding to each of the spatial nodes to obtain target communication flow prediction values ​​of multiple spatial nodes.

[0014] To achieve the above-mentioned purpose, the third aspect of an embodiment of the present application proposes an electronic device, which includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, it implements the traffic prediction method in wireless communication as described in the first aspect.

[0015] To achieve the above-mentioned purpose, the fourth aspect of an embodiment of the present application proposes a storage medium, which is a computer-readable storage medium and stores a computer program. When the computer program is executed by a processor, it implements the traffic prediction method in wireless communication as described in the first aspect.

[0016] The embodiments of the present application propose a method, device, electronic device, and storage medium for traffic prediction in wireless communications. First, historical spatiotemporal traffic data of multiple spatial nodes in a target area are acquired. Then, based on the historical spatiotemporal traffic data of each spatial node, multi-scale time features of the spatial nodes are extracted. Next, spatial dependency modeling is performed on the multi-scale time features to obtain a node embedding vector corresponding to each spatial node. Then, based on the node embedding vector corresponding to each spatial node, a target expert model is determined for the corresponding spatial node from multiple preset candidate expert models. Finally, the multi-scale time features are predicted using the target expert model corresponding to each spatial node to obtain target communication traffic prediction values ​​for the multiple spatial nodes.

[0017] With respect to the problems in the background technology, the traffic prediction method in wireless communication provided by the embodiment of the present application first extracts the multi-scale time features of the historical spatiotemporal traffic data, providing a feature basis that simultaneously contains multiple time-scale patterns for subsequent processing; then, by performing spatial dependency modeling on the multi-scale time features, it is possible to learn and generate a node embedding vector that can characterize the inherent attributes of each spatial node. Next, based on the node embedding vector, a target expert model is determined for each spatial node. This method can assign nodes with different traffic pattern characteristics to different expert models for processing according to the inherent attributes of the spatial node, thereby realizing adaptive modeling of spatial node heterogeneity. Finally, the multi-scale time features are predicted by the target expert model determined by its corresponding node embedding vector, and the prediction task can be completed using the expert model that is most suitable for processing a specific pattern. In this way, the present scheme can effectively learn the multi-scale time features in long-term series, and accurately model the complex heterogeneous interaction relationship between the multi-scale time features between different spatial nodes, thereby improving the prediction accuracy of spatiotemporal communication traffic.

[0018] Other features and advantages of the present application will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the present application. The purposes and other advantages of the present application can be achieved and obtained through the structures particularly pointed out in the description, claims and drawings. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 A flow chart of a method for traffic prediction in wireless communications provided in one embodiment of the present application; Figure 2 A flow chart of a method for traffic prediction in wireless communications provided in another embodiment of the present application; Figure 3 A flow chart of a method for traffic prediction in wireless communications provided in another embodiment of the present application; Figure 4A flow chart of a method for traffic prediction in wireless communications provided in another embodiment of the present application; Figure 5 A flow chart of a method for traffic prediction in wireless communications provided in another embodiment of the present application; Figure 6 A flow chart of a method for traffic prediction in wireless communications provided in another embodiment of the present application; Figure 7 A flow chart of a method for traffic prediction in wireless communications provided in another embodiment of the present application; Figure 8 A schematic diagram of a traffic prediction device in wireless communication provided by an embodiment of the present application; Figure 9 A schematic diagram of the hardware structure of an electronic device provided in yet another embodiment of the present application. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0021] It should be noted that although the functional modules are divided in the device schematic and the logical order is shown in the flowchart, in some cases, the steps shown or described can be performed in a different order than the module division in the device or the order in the flowchart.

[0022] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application pertains. The terms used herein are for the purpose of describing the embodiments of this application only and are not intended to limit this application.

[0023] Spatiotemporal forecasting of communication traffic is a key task in intelligent network management and resource scheduling. Its goal is to accurately predict future changes in communication traffic at different geographic locations based on historical data. With the development of 5G and future networks, network traffic exhibits significant periodicity and burstiness in the temporal dimension, and strong regional coupling and topological dependence in the spatial dimension. Therefore, effectively modeling the temporal dependencies and spatial correlations in communication traffic to achieve high-precision spatiotemporal traffic forecasting has become a critical issue for improving network efficiency and user service quality.

[0024] Graph neural networks are widely used in spatial modeling, such as spatiotemporal graph convolutional networks (GCNNs) that leverage predefined graphs or adaptive GCNs that can adaptively learn graph structures. For temporal modeling, the Mamba architecture, based on state-space models, has attracted considerable attention due to its high efficiency. However, these existing technologies suffer from two key drawbacks when processing long-term spatiotemporal data: first, the multi-scale information inherent in long-term sequences is difficult to learn simultaneously and effectively; second, the complex and heterogeneous interactions between these multi-scale temporal patterns across different spatial nodes cannot be fully modeled.

[0025] In response to the above problems, the traffic prediction method in wireless communication provided by the embodiment of the present application first extracts the multi-scale time features of historical spatiotemporal traffic data, providing a feature basis that simultaneously contains multiple time-scale patterns for subsequent processing; then, by performing spatial dependency modeling on the multi-scale time features, it is possible to learn and generate a node embedding vector that can characterize the inherent attributes of each spatial node. Next, based on the node embedding vector, a target expert model is determined for each spatial node. This method can assign nodes with different traffic pattern characteristics to different expert models for processing according to the inherent attributes of the spatial node, thereby realizing adaptive modeling of spatial node heterogeneity. Finally, the multi-scale time features are predicted by the target expert model determined by its corresponding node embedding vector, and the prediction task can be completed using the expert model that is most suitable for processing a specific pattern. In this way, the present scheme can effectively learn the multi-scale time features in long-term series, and accurately model the complex heterogeneous interaction relationship between the multi-scale time features between different spatial nodes, thereby improving the prediction accuracy of spatiotemporal communication traffic.

[0026] The following further describes a method, apparatus, electronic device, and storage medium for traffic prediction in wireless communication provided by embodiments of the present application. First, a method for traffic prediction in wireless communication in embodiments of the present application is specifically described.

[0027] Reference Figure 1 , which is an optional flowchart of the method for traffic prediction in wireless communication provided in an embodiment of the present application, Figure 1 The method may include but is not limited to steps 101 to 105. It is also understood that this embodiment is Figure 1 The order of step 101 to step 105 is not specifically limited, and the order of steps can be adjusted or some steps can be reduced or added according to actual needs.

[0028] Step 101: Obtain historical spatiotemporal traffic data of multiple spatial nodes in a target area.

[0029] Step 102: Based on the historical spatiotemporal traffic data of each spatial node, extract the multi-scale temporal features of the spatial node.

[0030] Step 103: Perform spatial dependency modeling on the multi-scale temporal features to obtain a node embedding vector corresponding to each spatial node.

[0031] Step 104: Based on the node embedding vector corresponding to each spatial node, determine a target expert model for the corresponding spatial node from a plurality of preset candidate expert models.

[0032] Step 105: predict the multi-scale time characteristics through the target expert model corresponding to each spatial node to obtain target communication flow prediction values ​​of multiple spatial nodes.

[0033] In step 101 of some embodiments, a "spatial node" can be a basic spatial unit after dividing the target area, such as a geographic grid or a base station coverage area. "Historical spatiotemporal traffic data" refers to the communication traffic information recorded at these spatial nodes over a continuous period of time. This data provides the necessary input for subsequent models to learn the spatiotemporal evolution of traffic.

[0034] See also Figure 2 In some embodiments, step 101 may include, but is not limited to, steps 201 to 203 .

[0035] Step 201: Obtain device measurement data reported by all terminal devices in the target area.

[0036] Step 202: Divide the target area into a plurality of preset geographic grids, and determine each geographic grid as a spatial node.

[0037] Step 203: Determine the historical traffic data of each spatial node in each preset time segment based on the geographic location data and the traffic data, and obtain the historical spatiotemporal traffic data of each spatial node.

[0038] In step 201 of some embodiments, device measurement data reported by all terminal devices within the target coverage area of ​​the communication network is first obtained from edge nodes (such as base stations) of the communication network. "Terminal devices" here generally refer to user equipment (UE), and the "device measurement data" reported by them technically corresponds to user measurement reports (MRs). This raw data is high-frequency and massive, containing key information necessary for subsequent processing. Specifically, it includes geographic location data that can identify the device's location, such as longitude and latitude coordinates, and traffic data that can reflect network usage, such as downlink throughput. This step is the source of data for the entire prediction process.

[0039] In step 202 of some embodiments, in order to convert the continuous geographic space into a discrete structure that can be processed by the model, the target area needs to be spatially divided. Specifically, this step divides the map data of the entire target area into multiple preset geographic grids of uniform size and non-overlapping. For example, the area can be divided into 10m x 10m grids. After the rasterization is completed, this embodiment logically determines each independent geographic grid as a "spatial node", which provides a basis for the subsequent use of graph neural networks and other methods for spatial dependency modeling, so that complex geographic spatial relationships can be abstracted into a graph structure composed of nodes and edges.

[0040] In step 203 of some embodiments, after the discretization of the spatial dimension is completed, the time dimension needs to be discretized, that is, the continuous time is divided into multiple preset time segments, for example, with a time granularity of 5 minutes. Then, based on the geographic location data and traffic data obtained in step 201, each device measurement data is attributed to the specific geographic grid in which it falls within the specific time segment. By fusing the traffic data of all terminal devices falling in the same geographic grid and the same time segment, the historical traffic data of the spatial node in the time segment can be determined. Traverse all spatial nodes and all historical time segments and repeat this process, and the data finally obtained is the historical spatiotemporal traffic data of each spatial node in a regular format that can be directly used by subsequent models.

[0041] Through the above steps 201 to 203, this embodiment can convert massive, disordered original measurement data into a structured, discrete spatiotemporal data format. Specifically, through spatial rasterization and time slicing, a unified data view is provided for subsequent spatiotemporal modeling. Determining the geographic grid as a spatial node allows the continuous geographic space to be represented by a discrete graph structure, which is the basis for performing spatial modeling operations such as graph convolution. Ultimately, by performing traffic statistics and aggregation on each spatiotemporal unit, data redundancy and noise are effectively reduced, and the key information of the traffic dynamics in a specific area within a specific time period can be reflected macroscopically, thereby providing high-quality, uniformly formatted data input for subsequent models to perform high-precision spatiotemporal traffic predictions.

[0042] In step 102 of some embodiments, because communication traffic exhibits both long-term periodicity (e.g., daily or weekly variations) and short-term bursts (e.g., minute-level fluctuations), this step aims to process the raw historical data to explicitly represent these temporal patterns at different scales at the feature level. The resulting "multi-scale temporal features" are a richer and more comprehensive representation that can simultaneously characterize both long-term trends and short-term variations in traffic.

[0043] See also Figure 3In some embodiments, step 102 may include, but is not limited to, steps 301 to 303 .

[0044] Step 301: Project the historical spatiotemporal traffic data of each spatial node into a latent space feature tensor through a multi-layer perceptron; wherein the latent space feature tensor includes a preset latent space feature dimension.

[0045] Step 302: Convolution processing is performed on each latent space feature tensor through multiple one-dimensional convolution layers with different preset convolution kernel sizes to obtain time feature representations at multiple time scales.

[0046] Step 303: For each spatial node, stack the temporal feature representations at multiple time scales to obtain a multi-scale temporal feature of the spatial node.

[0047] In step 301 of some embodiments, after obtaining the input tensor represented as After the historical spatiotemporal traffic data is processed, it is first processed by an input mapping layer composed of a multi-layer perceptron (MLP). The multi-layer perceptron here is a neural network that transforms the input tensor into a Converted into a high-dimensional latent space feature tensor , where T is the historical time step, N is the number of spatial nodes, C is the input feature dimension, and d is the preset, usually higher-dimensional latent space feature dimension.

[0048] In step 302 of some embodiments, in order to decompose patterns of different scales from the time dimension, this embodiment uses multiple (e.g., q) preset one-dimensional convolution layers with different convolution kernel sizes. , perform parallel convolution processing on the latent space feature tensor H obtained in the previous step. For each scale q, the corresponding convolution operation can be expressed by the following formula:

[0049] in, Represents the convolution kernel size used by the qth convolution layer. Different convolution kernel sizes have different temporal receptive fields. For example, a larger Able to capture long-term trends, while smaller By executing the convolution operation represented by this formula q times in parallel, we can obtain q groups of time feature representations representing different time scales. .

[0050] In step 303 of some embodiments, after obtaining multiple groups of time feature representations corresponding to different time scales, this step integrates these feature representations. Specifically, this step integrates the q groups of time feature representations obtained in step 302 Stacking along a new dimension forms a higher-dimensional, uniformly structured tensor. This final tensor is the "multi-scale temporal feature" of the spatial node, and its dimension can be expressed as This operation completely preserves and integrates all decomposed and extracted temporal information at different scales into a unified data structure, so that subsequent model layers can simultaneously obtain and utilize temporal information at all scales.

[0051] Through the above steps 301 to 303, this embodiment can effectively convert the original spatiotemporal traffic data into a deep feature representation that is more conducive to model learning and includes multi-scale time dimensions. Specifically, the multi-layer perceptron projection in step 301 elevates the data to a high-dimensional latent space, thereby enhancing the expressive power of the features. Step 302 achieves effective decomposition of the time series by using multiple one-dimensional convolution kernels of different sizes. It can simultaneously capture and separate multiple patterns present in the data, such as long-term regularities and short-term sudden fluctuations. Finally, step 303, through a stacking operation, completely retains and integrates these separated information at different scales into a unified data structure. This series of operations provides high-quality feature inputs that have decoupled information at different time scales for subsequent spatial modeling and prediction steps, thereby solving the problem of difficulty in simultaneously learning multi-scale features in the background technology.

[0052] In some embodiments, step 103 aims to learn and quantify the relationships and inherent properties between different spatial nodes. Spatial dependency modeling can analyze and capture the complex network relationships that influence how traffic changes at one node are affected by other nodes. The resulting "node embedding vector" is a numerical vector that represents the unique static characteristics of each spatial node, encompassing the node's geographic or functional attributes.

[0053] See also Figure 4 In some embodiments, step 103 may include, but is not limited to, steps 401 to 403 .

[0054] Step 401: Based on the preset initial node embedding vector of each spatial node, an adaptive graph adjacency matrix is ​​generated for representing the spatial dependency relationship between multiple spatial nodes.

[0055] Step 402: Based on the adaptive graph adjacency matrix, perform graph convolution processing on the multi-scale temporal features to aggregate spatial neighborhood information of multiple spatial nodes to obtain multi-scale temporal aggregated features.

[0056] Step 403: Perform cross-scale fusion on the multi-scale temporal aggregation features to obtain a node embedding vector corresponding to each spatial node.

[0057] In step 401 of some embodiments, in order to model the interdependence between spatial nodes, the method first generates an adaptive graph adjacency matrix. The core of this step is that it does not rely on any predefined, fixed geographic topology information, but allows the model to learn the strength of the association between nodes from the data. Specifically, this step presets a learnable "initial node embedding vector" for each spatial node. This set of vectors is represented in the model as a parameter matrix By calculating the similarity between these embedding vectors, we can generate an adaptive graph adjacency matrix A that can characterize the functional associations between nodes. In practical applications, its normalized form is usually used, which can be obtained by the following formula:

[0058] Where D is the degree matrix. This process enables the model to capture non-intuitive but real spatial dependencies.

[0059] In step 402 of some embodiments, after constructing the adaptive graph adjacency matrix, graph convolution processing can be used to aggregate spatial information. This step applies the spatial relationship represented by the adaptive graph adjacency matrix generated in the previous step to the "multi-scale temporal features". For each time scale, the graph convolution operation enables each spatial node to aggregate the feature information of its neighboring nodes in the adaptive graph, thereby integrating the node representation into the context of its spatial environment. A specific implementation of this operation can be approximately represented by the following formula:

[0060] in, is the specific scale feature of the input, It is the output after spatial information aggregation, namely "multi-scale time aggregation feature". is a low-rank parameter used to reduce computational complexity.

[0061] In step 403 of some embodiments, after completing spatial information aggregation for temporal features of different scales respectively, cross-scale fusion is also performed on the multi-scale temporal aggregation features. The purpose of this step is to effectively integrate the feature information separated at different scales and already containing spatial context. A fusion rule can be designed to ensure that long-period, coarse-grained features can effectively guide short-period, fine-grained features, thereby avoiding mutual interference between information of different scales. After processing this cross-scale fusion step, the model finally learns and determines the "node embedding vector" that can accurately characterize each spatial node in the multi-dimensional spatiotemporal context.

[0062] See also Figure 5 In some embodiments, step 403 may include, but is not limited to, steps 501 to 505 .

[0063] Step 501: Project the multi-scale temporal aggregation features into a query matrix, a key matrix, and a value matrix respectively.

[0064] Step 502: Perform matrix multiplication on the query matrix and the key matrix, and scale and mask the multiplication result to obtain an attention score.

[0065] Step 503: Normalize the attention score and perform matrix multiplication with the value matrix to obtain the attention output feature.

[0066] Step 504: Perform residual connection and layer normalization on the attention output features and the multi-scale temporal aggregation features to generate fused spatiotemporal features.

[0067] Step 505: Based on the fused spatiotemporal features, determine the node embedding vector corresponding to each spatial node.

[0068] In step 501 of some embodiments, in order to perform attention calculation, it is first necessary to input multi-scale temporal aggregation features ( ) is linearly transformed. Specifically, through three linear projection layers, the multi-scale temporal aggregation features are converted into three different matrices, namely the query matrix (Query, Q), the key matrix (Key, K), and the value matrix (Value, V). The transformation process can be expressed as:

[0069] This operation maps the original features into different subspaces, laying the foundation for the subsequent calculation of inter-feature correlations (through Q and K) and extraction of feature content (through V).

[0070] In step 502 of some embodiments, after obtaining the query, key, and value matrices, this step calculates the attention scores by computing their relationships. This process first performs matrix multiplication on the query matrix Q and the key matrix K to calculate the raw correlation scores between different scale features. Then, for the stability of the training process, the result is usually scaled, for example, divided by a scaling factor. . The most crucial operation in this step is the masking process, that is, adding a preset "lower triangular mask" (M) to the above result. The specific implementation of this mask matrix M can be that when the scale index i < j, the corresponding value is negative infinity, and the other positions are 0. Mathematically, this makes the features at any scale i unable to focus on the features at a finer scale j than it, thus forcing a one-way information flow from coarse-grained features to fine-grained features. The process formula can be expressed as: In step 503 of some embodiments, in order to convert the attention scores obtained in the previous step into available weights and extract information based on them, this step first applies a normalization function, usually the Softmax function, to the masked attention scores. This operation converts the scores into a probability distribution with a sum of 1 along a specific dimension, that is, the attention weights, which quantitatively represent how much attention the model should pay to the features at which scales during the fusion process. Subsequently, matrix multiplication is performed on this normalized attention weight and the value matrix V. This operation is functionally equivalent to a weighted sum, which selectively extracts and aggregates information from the value matrix V according to the attention weights to obtain the "attention output feature" ( ), and the above calculation formula can be expressed as:

[0071] In step 504 of some embodiments, in order to enable this deep attention module to be stably trained and avoid performance degradation in information transmission, this embodiment performs a residual connection (Residual Connection) on the "attention output feature" obtained in step 503 and the original input "multi-scale temporal aggregation feature" of this fusion process, so that information can directly propagate across layers, effectively alleviating the vanishing gradient problem in deep networks. Then, layer normalization is performed on the added result again to normalize the distribution of each layer of features, further stabilizing the training process and accelerating model convergence. The final output of this step is the "fused spatio-temporal feature" ( ), and the formula can be expressed as:

[0072] In step 505 of some embodiments, based on the "fused spatiotemporal features" obtained in the previous step, which have deeply integrated multi-scale spatiotemporal information, this embodiment ultimately determines the "node embedding vector corresponding to each spatial node." The node embedding vector output in this step can be considered the final form of the initial node embedding vector after thorough learning and optimization throughout the spatial dependency modeling and cross-scale fusion process.

[0073] By executing the above steps 501 to 505, this embodiment uses QKV projection and attention calculation, and the model can adaptively learn the mutual importance between features of different time scales. The core lies in the application of the lower triangle mask in step 502, which forces the implementation of a unidirectional information flow from coarse to fine, ensuring that long-term regularities can serve as effective prior knowledge to guide the understanding of short-term time series patterns, while avoiding information leakage and interference, and solving the problem of disordered interaction of multi-scale information. In addition, the introduction of residual connections and layer normalization in step 504 enables multi-feature fusion. The final node embedding vector is a node representation associated with multi-scale spatiotemporal contexts, which greatly improves the discriminability and effectiveness of features.

[0074] In step 104 of some embodiments, this embodiment employs a hybrid expert approach, presetting a set of "candidate expert models" with varying processing capabilities. This step is an intelligent routing or selection process that uses the "node embedding vectors" obtained in the previous step, which characterize node characteristics, as a basis for assigning one or more "target expert models" to each spatial node that are most suitable for processing its data pattern. This step enables the method to adaptively process spatial nodes with varying characteristics.

[0075] See also Figure 6 In some embodiments, the plurality of preset candidate expert models are set to include a shared expert model and a plurality of dedicated expert models, and step 104 may include, but is not limited to, steps 601 to 602.

[0076] Step 601: Calculate the routing score corresponding to each dedicated expert model based on the node embedding vector.

[0077] Step 602: The dedicated expert model and the shared expert model with the highest selection routing score are jointly determined as the target expert model.

[0078] In step 601 of some embodiments, in order to match the most appropriate dedicated expert model to each spatial node, the present embodiment first needs to calculate the routing score of each dedicated expert model for the node. The basis for calculating the routing score is not the dynamically changing time series itself, but the "node embedding vector" (which can be expressed as) obtained in the previous step (such as step 103) and capable of representing the static inherent properties of the spatial node. ). Specifically, a gating network transforms the node’s As input, it outputs a scalar value for each dedicated expert model to be selected (for example, there are K in total), which is the "routing score" corresponding to the dedicated expert. For the kth dedicated expert, the calculation of its routing score can be expressed as ,in is the corresponding gating function. In some specific embodiments, in order to enhance the robustness of the routing process, a small random noise rk may be added to the calculated score.

[0079] In some embodiments, in step 602, after calculating the routing scores of all dedicated expert models for the current spatial node, this step selects and ultimately determines the target expert model for the node based on the scores. The specific selection strategy is to select the dedicated expert model with the highest routing score. This selection process can be expressed as follows:

[0080] Here, g is the index of the selected dedicated expert model. The "target expert model" ultimately determined in this step is a combination of two parts: the dedicated expert model just selected by the routing score as the most suitable for processing the node's unique pattern; and the "shared expert model" used by all spatial nodes. Subsequent prediction steps will utilize both expert models simultaneously.

[0081] By executing steps 601 to 602 above, this embodiment details an adaptive routing based on a Mixture of Experts (MoE) architecture. This architecture does not make routing decisions based on dynamic, time-varying time series data itself. Instead, it utilizes the node embedding vectors input in step 601, which characterize the static, inherent properties of spatial nodes, as the routing basis. This makes routing decisions smoother and more stable, avoiding the frequent jumps in expert selection caused by minor fluctuations in the input sequence. The selection mechanism in step 602 ensures that the data of each spatial node is refined by a dedicated expert who is most adept at processing its specific patterns. At the same time, the presence of shared experts ensures that all nodes have access to common, shared pattern information. This combination of shared and dedicated approaches enhances the model's expressiveness and generalization capabilities, effectively addressing the difficulty in modeling the complex, heterogeneous interactions between different spatial nodes, a problem previously encountered in prior art.

[0082] In step 105 of some embodiments, the multi-scale temporal features extracted in the previous step are fed into the "target expert model" determined for the node in step 104. Because these expert models are tailored to the node's characteristics, they can more specifically analyze and infer temporal patterns in the input features. Ultimately, the model outputs a prediction of the communication traffic volume for one or more future time steps, namely the "target communication traffic volume prediction value."

[0083] In some embodiments, step 105 may include, but is not limited to: The multi-scale temporal features are extracted using the selected dedicated expert model and the shared expert model to obtain dedicated features and shared features, respectively. The feature extraction process includes: performing time scale expansion processing on the multi-scale temporal features through multiple one-dimensional convolutional layers with different convolution kernel sizes to obtain expanded multi-scale temporal features; Perform multi-scale feature extraction on the expanded multi-scale temporal features through a preset learnable multi-scale bias; Fuse the dedicated features and shared features to obtain the target fused features; The target fusion features are mapped into prediction space through a multi-layer perceptron to obtain the target communication flow prediction value.

[0084] First, this method uses a selected dedicated expert model and a shared expert model to perform parallel feature extraction on the "multi-scale temporal features" obtained in the previous step. The input feature data stream is replicated and fed simultaneously into two expert models with identical structures but independent parameters. After a series of internal processing, these two expert models each output two different features: a "dedicated feature" (generated by the dedicated expert model) that reflects the unique patterns of that spatial node, and a "shared feature" (generated by the shared expert model) that reflects the common patterns across all spatial nodes.

[0085] In the above feature extraction process, each expert model performs a set of identical processing processes. The processing first includes a time scale expansion process, which processes the input multi-scale time features again through a set of one-dimensional convolution layers with different convolution kernel sizes to further expand the receptive field of the model and enable it to capture longer-term dependencies. Subsequently, the expanded multi-scale time features are subjected to multi-scale feature extraction processing, using a sequence processing method based on the state space model (SSM), and by adding internal parameters of the method (specifically, the selective parameters ) introduces a “preset learnable multi-scale bias” (Ω), which can be influenced by the following formula:

[0086] This bias enables a single processing model to effectively distinguish and identify features of different scales that are spliced ​​together, and to implement independent and targeted pattern extraction for them, thus solving the problem that traditional sequence models have difficulty in processing multi-scale information.

[0087] After obtaining the "dedicated features" and "shared features" respectively, this method fuses these two features to obtain a "target fused feature" that contains both common information and characteristic information. In some specific embodiments, this feature fusion can be achieved through a simple element-by-element addition operation, and the process can be expressed as follows:

[0088] in, represents the output of a dedicated expert model, represents the output of the shared expert model, and z is the "target fusion feature".

[0089] Finally, after obtaining the target fusion features, this method maps them into the prediction space through an output mapping layer consisting of a multi-layer perceptron (MLP). This step transforms the high-dimensional, abstract fusion features into a final output with clear physical meaning. The output mapping layer maps the fusion feature dimensions to the same dimensions as the prediction target, ultimately obtaining the target traffic flow prediction value for one or more future time steps.

[0090] By executing the above steps, this embodiment, through the combination of dedicated experts and shared experts, can simultaneously capture the unique patterns of specific spatial nodes and the general laws shared by all nodes, greatly enhancing the modeling capabilities of spatially heterogeneous data. Its core advantage lies in the feature extraction process within the expert model: this process effectively solves the problem that traditional sequence models are difficult to process multi-scale information simultaneously by expanding the time scale and introducing learnable multi-scale bias processing. This multi-scale bias enables the model to perform independent and targeted analysis and extraction of features at different time scales within a unified framework. Ultimately, through feature fusion and output mapping, these deeply and finely processed features are transformed into high-precision prediction results.

[0091] See also Figure 7 In some embodiments, after step 105, the process may also include, but not be limited to, steps 701 to 702.

[0092] Step 701: Based on a preset joint loss function, calculate the loss value between the target communication traffic prediction value and the corresponding actual traffic value.

[0093] Step 702: Based on the loss value, update multiple candidate expert models.

[0094] In step 701 of some embodiments, after the model generates the "target communication traffic prediction value," in order to quantitatively evaluate the model's performance and guide its parameter optimization, this step calculates the loss between the prediction value and the corresponding actual traffic value based on a preset joint loss function. This joint loss function is an important component of the present invention and consists of two key parts, which can be expressed by the following total loss function formula:

[0095] Where L is the total loss and λ is a hyperparameter that balances the importance of the two parts of the loss. The following is a detailed description of the two components of the joint loss function. The first part is the mean absolute error loss, expressed as Its main function is to directly measure the numerical difference between the predicted value and the true value. As the core indicator of the model prediction accuracy, the formula is:

[0096] The second part is the causal contrast loss, expressed as , which is designed to promote pattern decoupling between multiple candidate expert models and encourage different specialized experts to learn different patterns. To achieve this goal, the loss first defines a special "asymmetric similarity measure":

[0097] Among them, γ1 and γ2 are two adjustable hyperparameters used to control the decay rate of the similarity score with the change of scale difference. Based on this similarity metric, the causal contrast loss can be calculated by the following formula, which aims to bring the similarity of positive sample pairs closer and push the similarity of negative sample pairs further away:

[0098] in, Represents the set of data samples that are routed to the k-th expert in a training batch. Represents "anchor samples", that is, samples from the set The feature representation of the i-th sample in the p-th scale is, Represents the set of "positive samples" corresponding to the above "anchor samples". In contrastive learning, positive samples usually refer to samples that should be considered similar to anchor samples in terms of semantics or patterns. Represents the set of “negative samples” corresponding to the “anchor samples”, that is, samples that should be considered dissimilar in pattern. is a positive sample taken from the positive sample set, is a negative sample taken from the negative sample set, and θ is the temperature hyperparameter.

[0099] In step 702 of some embodiments, after calculating the total loss value L, this step will adjust and optimize all learnable parameters in the model based on the loss value. The ultimate goal of this step is to minimize the loss value. In the practice of deep learning, this process is usually achieved through a gradient optimization algorithm (such as Adam, SGD, etc.) combined with a backpropagation mechanism. The loss value is used as a starting point, and its gradient is backpropagated to each learnable parameter of the model, including the weights of all convolutional layers and linear layers, the learnable bias in the multi-scale Mamba, and the initial node embedding vector in the adaptive graph network. By iteratively executing this step, all parameters of the model will be fine-tuned in the direction of continuously reducing the total loss, so that the model reaches the optimal state in terms of both prediction accuracy and internal expert function differentiation.

[0100] By executing the above steps 701 to 702, this embodiment ensures that the core goal of model training is to continuously improve the accuracy of predictions and make the predicted values ​​converge to the true values ​​through the mean absolute error loss in step 701. On the other hand, by introducing the innovative causal contrast loss, this method can guide different specialized experts to learn independent and different data patterns from the training mechanism. It effectively avoids the problem of expert function redundancy and ensures that each expert can become a model for processing specific patterns, thereby maximizing the advantages of the hybrid expert architecture in dealing with complex heterogeneous data. Finally, step 702 optimizes the model as a whole based on the joint loss value to improve the overall prediction accuracy.

[0101] See also Figure 8 The embodiment of the present application further provides a traffic prediction device in wireless communication, which can implement the above-mentioned traffic prediction method in wireless communication, including: An acquisition module is used to obtain historical spatiotemporal traffic data of multiple spatial nodes in the target area; An extraction module is used to extract multi-scale temporal features of spatial nodes based on the historical spatiotemporal traffic data of each spatial node; The modeling module is used to perform spatial dependency modeling on multi-scale temporal features and obtain the node embedding vector corresponding to each spatial node; a determination module, configured to determine a target expert model from a plurality of preset candidate expert models for the corresponding spatial node based on a node embedding vector corresponding to each spatial node; The prediction module is used to predict the multi-scale time characteristics through the target expert model corresponding to each spatial node to obtain the target communication flow prediction values ​​of multiple spatial nodes.

[0102] The embodiments of the present application propose a method, device, electronic device, and storage medium for traffic prediction in wireless communications. First, historical spatiotemporal traffic data of multiple spatial nodes in a target area are acquired. Then, based on the historical spatiotemporal traffic data of each spatial node, multi-scale time features of the spatial nodes are extracted. Next, spatial dependency modeling is performed on the multi-scale time features to obtain a node embedding vector corresponding to each spatial node. Then, based on the node embedding vector corresponding to each spatial node, a target expert model is determined for the corresponding spatial node from multiple preset candidate expert models. Finally, the multi-scale time features are predicted using the target expert model corresponding to each spatial node to obtain target communication traffic prediction values ​​for the multiple spatial nodes.

[0103] With respect to the problems in the background technology, the traffic prediction method in wireless communication provided by the embodiment of the present application first extracts the multi-scale time features of the historical spatiotemporal traffic data, providing a feature basis that simultaneously contains multiple time-scale patterns for subsequent processing; then, by performing spatial dependency modeling on the multi-scale time features, it is possible to learn and generate a node embedding vector that can characterize the inherent attributes of each spatial node. Next, based on the node embedding vector, a target expert model is determined for each spatial node. This method can assign nodes with different traffic pattern characteristics to different expert models for processing according to the inherent attributes of the spatial node, thereby realizing adaptive modeling of spatial node heterogeneity. Finally, the multi-scale time features are predicted by the target expert model determined by its corresponding node embedding vector, and the prediction task can be completed using the expert model that is most suitable for processing a specific pattern. In this way, the present scheme can effectively learn the multi-scale time features in long-term series, and accurately model the complex heterogeneous interaction relationship between the multi-scale time features between different spatial nodes, thereby improving the prediction accuracy of spatiotemporal communication traffic.

[0104] An embodiment of the present application further provides an electronic device, including: at least one memory; at least one processor; at least one program; The program is stored in the memory, and the processor executes the at least one program to implement the traffic prediction method in wireless communication implemented in the present application. The electronic device can be any smart terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), an in-vehicle computer, etc.

[0105] See also Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is shown. The electronic device includes: The processor 901 may be implemented as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute relevant programs to implement the technical solutions provided in the embodiments of the present application. The memory 902 can be implemented in the form of ROM (Read Only Memory), static storage device, dynamic storage device or RAM (Random Access Memory). The memory 902 can store an operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called by the processor 901 to execute the traffic prediction method in wireless communication in the embodiments of this application. Input / output interface 903, used to implement information input and output; Communication interface 904, used to implement communication interaction between this device and other devices, which can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WiFi, Bluetooth, etc.); Bus 905 , which transmits information between various components of the device (e.g., processor 901 , memory 902 , input / output interface 903 , and communication interface 904 ); The processor 901 , the memory 902 , the input / output interface 903 and the communication interface 904 are connected to each other in communication within the device via a bus 905 .

[0106] An embodiment of the present application further provides a storage medium, which is a computer-readable storage medium and stores a computer program. When the computer program is executed by a processor, the above-mentioned traffic prediction method in wireless communication is implemented.

[0107] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely arranged relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0108] The embodiments described in the embodiments of this application are intended to more clearly illustrate the technical solutions of the embodiments of this application and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0109] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.

[0110] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, i.e., they may be located in one place or distributed across multiple network units. Some or all of the modules may be selected based on actual needs to achieve the objectives of this embodiment.

[0111] Those skilled in the art will appreciate that all or some of the steps in the methods, systems, and functional modules / units in the devices disclosed above may be implemented as software, firmware, hardware, or appropriate combinations thereof.

[0112] The terms "first," "second," "third," "fourth," and the like (if any) in the specification of the present application and the accompanying drawings are used to distinguish similar objects and are not necessarily used to describe a particular order or precedence. It should be understood that the terms used in this manner are interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in orders other than those illustrated or described herein. In addition, the terms "including" and "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such process, method, product, or apparatus.

[0113] It should be understood that in this application, "at least one (item)" means one or more, and "more" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c A), can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0114] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the above-mentioned units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0115] The units described above as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0116] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0117] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store programs.

[0118] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but are not intended to limit the scope of the present invention. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and essence of the present invention should be within the scope of the present invention.

Claims

1. A method for traffic prediction in wireless communication, characterized in that: The method comprises: Obtain historical spatiotemporal traffic data for multiple spatial nodes in the target area; Extracting multi-scale time features of each spatial node based on the historical spatiotemporal traffic data of the spatial node; Performing spatial dependency modeling on the multi-scale temporal features to obtain a node embedding vector corresponding to each of the spatial nodes; Based on the node embedding vector corresponding to each of the spatial nodes, determining a target expert model for the corresponding spatial node from a plurality of preset candidate expert models; The multi-scale time characteristics are predicted by the target expert model corresponding to each of the spatial nodes to obtain target communication flow prediction values ​​of multiple spatial nodes.

2. The method according to claim 1, characterized in that The extracting of multi-scale time features of each spatial node based on the historical spatiotemporal traffic data of the spatial node includes: Projecting the historical spatiotemporal traffic data of each spatial node into a latent space feature tensor through a multi-layer perceptron; wherein the latent space feature tensor includes a preset latent space feature dimension; Performing convolution processing on each of the latent space feature tensors through a plurality of preset one-dimensional convolution layers with different convolution kernel sizes to obtain temporal feature representations at multiple time scales; For each of the spatial nodes, the time feature representations at multiple time scales are stacked to obtain the multi-scale time feature of the spatial node.

3. The method according to claim 1, characterized in that The acquisition of historical spatiotemporal traffic data of multiple spatial nodes in the target area includes: Obtaining device measurement data reported by all terminal devices in the target area; wherein the device measurement data includes geographic location data and traffic data; Dividing the target area into a plurality of preset geographic grids, and determining each of the geographic grids as a spatial node; Based on the geographic location data and the traffic data, the historical traffic data of each of the spatial nodes in each preset time segment is determined to obtain the historical spatiotemporal traffic data of each of the spatial nodes.

4. The method according to claim 1, wherein The performing spatial dependency modeling on the multi-scale temporal features to obtain a node embedding vector corresponding to each of the spatial nodes includes: generating an adaptive graph adjacency matrix for characterizing spatial dependency relationships between the plurality of spatial nodes based on a preset initial node embedding vector for each of the spatial nodes; Based on the adaptive graph adjacency matrix, performing graph convolution processing on the multi-scale temporal features to aggregate spatial neighborhood information of the multiple spatial nodes to obtain multi-scale temporal aggregated features; The multi-scale temporal aggregation features are cross-scale fused to obtain a node embedding vector corresponding to each of the spatial nodes.

5. The method according to claim 4, characterized in that The cross-scale fusing of the multi-scale temporal aggregation features to obtain a node embedding vector corresponding to each of the spatial nodes includes: Projecting the multi-scale temporal aggregation features into a query matrix, a key matrix, and a value matrix respectively; Performing matrix multiplication on the query matrix and the key matrix, and scaling and masking the multiplication result to obtain an attention score; wherein the masking adopts a preset lower triangular mask; Normalizing the attention score and performing matrix multiplication with the value matrix to obtain an attention output feature; Performing residual connection and layer normalization processing on the attention output feature and the multi-scale temporal aggregation feature to generate a fused spatiotemporal feature; Based on the fused spatiotemporal features, the node embedding vector corresponding to each of the spatial nodes is determined.

6. The method according to claim 1, characterized in that The plurality of preset candidate expert models are set to include a shared expert model and a plurality of dedicated expert models, and the determining of a target expert model from the plurality of preset candidate expert models for the corresponding spatial node based on the node embedding vector corresponding to each spatial node includes: Calculating a routing score corresponding to each of the dedicated expert models based on the node embedding vector; The dedicated expert model and the shared expert model with the highest selection routing score are jointly determined as the target expert model.

7. The method according to claim 6, characterized in that The predicting the multi-scale time feature by the target expert model corresponding to each of the spatial nodes to obtain target communication flow prediction values ​​of the plurality of spatial nodes includes: The multi-scale time features are respectively extracted using the selected dedicated expert model and the shared expert model to obtain dedicated features and shared features; wherein the feature extraction process includes: Performing time scale expansion processing on the multi-scale time feature through multiple one-dimensional convolution layers with different convolution kernel sizes to obtain an expanded multi-scale time feature; Performing multi-scale feature extraction processing on the expanded multi-scale temporal features through a preset learnable multi-scale bias; Fusing the dedicated features with the shared features to obtain target fused features; The target fusion features are mapped into a prediction space by a multi-layer perceptron to obtain the target communication flow prediction value.

8. The method according to claim 1, characterized in that After obtaining the target communication flow prediction values ​​of the plurality of spatial nodes, the method further includes: Based on a preset joint loss function, the loss value between the target communication traffic prediction value and the corresponding actual traffic value is calculated; wherein the joint loss function includes a mean absolute error loss and a causal contrast loss, the mean absolute error loss is used to measure the prediction error, and the causal contrast loss is used to promote pattern decoupling between multiple candidate expert models; Based on the loss value, multiple candidate expert models are updated.

9. A traffic prediction device in wireless communication, characterized in that: include: An acquisition module is used to obtain historical spatiotemporal traffic data of multiple spatial nodes in the target area; an extraction module, configured to extract multi-scale time features of each spatial node based on the historical spatiotemporal traffic data of the spatial node; A modeling module, configured to perform spatial dependency modeling on the multi-scale temporal features to obtain a node embedding vector corresponding to each of the spatial nodes; a determination module, configured to determine a target expert model from a plurality of preset candidate expert models for the corresponding spatial node based on the node embedding vector corresponding to each spatial node; The prediction module is used to predict the multi-scale time characteristics through the target expert model corresponding to each of the spatial nodes to obtain target communication flow prediction values ​​of multiple spatial nodes.

10. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the method for traffic prediction in wireless communication according to any one of claims 1 to 8 is implemented.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method for traffic prediction in wireless communication according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Network traffic prediction method based on combination of GNN-LSTM

    CN112906982A

  • Network traffic prediction method and device, electronic equipment and storage medium

    CN118200167A

  • Traffic flow prediction method based on multi-scale joint space-time hypergraph neural network

    CN118966479A

  • Traffic determination method and apparatus based on spatio-temporal data, and device and medium

    WO2023207411A1