Base station flow prediction method and device
By using one-dimensional convolutional network, time self-attention length memory module and space multi-head attention module in the railway 5G private network base station traffic prediction, the problem that the existing technology cannot effectively capture the spatial and temporal changes of base station traffic is solved, and higher prediction accuracy and lower operating costs are achieved.
Patent Information
- Application Number
- CN202510024190.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-01-07
AI Technical Summary
The existing technology is not effective in predicting the traffic of the base station of the railway 5G private network, and it is impossible to effectively capture the spatial and temporal changes of the base station traffic.
A base station traffic prediction method is adopted. By obtaining the traffic spatiotemporal sequence and the base station traffic prediction model, a one-dimensional convolution network, a time self-attention length memory module and a spatial multi-head attention module are used to extract the spatial correlation between adjacent base station traffic and the temporal correlation of the base station traffic sequence, combining the self-attention mechanism and the multi-head attention mechanism to model sudden and periodic changes, and predict through feature fusion.
It improves the accuracy of base station traffic prediction, can more comprehensively capture the changes in base station traffic, handle complex relationships in spatio-temporal sequence data, reduces energy consumption and operation costs, and improves the stability of network services.
Smart Images

Figure CN120034873A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of base station traffic prediction, and particularly to a base station traffic prediction method and device. Background Art
[0002] With the construction requirements of intelligent transportation and the proposal of the intelligent high-speed rail system architecture, the ubiquitous interconnection requirements of intelligent high-speed rail pose a huge bandwidth demand for vehicle-ground wireless data transmission. At the same time, due to the high reliability requirements of the high-speed rail train control system for vehicle-ground wireless data transmission, it has given rise to the need to upgrade the railway wireless communication to the next-generation broadband communication system with higher transmission rates and to build a railway 5G private network. However, due to the increase in the number of railway 5G private network base stations and the obvious increase in the energy consumption of a single base station, the importance of base station energy-saving technology is increasing day by day. It is necessary to accurately guide the energy-saving control decision of the base station through base station traffic prediction.
[0003] Currently, the research on base station traffic prediction mainly focuses on the mathematical characteristics of service data itself, rather than the special application scenarios of services. In the spatial dimension, the railway 5G private network base stations are arranged in a one-dimensional line. In the time dimension, the trains run intermittently, there is a corresponding skylight period for maintenance at night, and the services have characteristics such as suddenness and periodicity. These all result in a large difference between the service traffic of railway 5G private network base stations and that in the civil scenario of public mobile communication. The effect of using the conventional public network base station traffic prediction method is not ideal. Summary of the Invention
[0004] The present invention provides a base station traffic prediction method and device to solve the defect that the effect of the conventional public network base station traffic prediction method is not ideal, which can capture the change law of base station traffic more comprehensively, process the complex relationship in spatio-temporal sequence data more effectively, and improve the prediction accuracy. The technical solutions proposed by the present invention are as follows: In the first aspect, the present invention provides a base station traffic prediction method, including: Obtaining a traffic spatio-temporal sequence and a base station traffic prediction model; wherein, the traffic spatio-temporal sequence includes a first traffic sequence and a second traffic sequence, and the base station traffic prediction model includes a first branch, a second branch, and an output layer; The first branch extracts the spatial correlation between adjacent base station traffic and the temporal correlation of a single base station traffic sequence based on the first traffic sequence, and uses a self-attention mechanism to model the sudden change of the same base station traffic in time in the time dimension, and uses a multi-head attention mechanism to model the sudden change distribution law of traffic between different base stations in the same period in the spatial dimension, so as to obtain a first feature tensor; The second branch extracts the temporal correlation of the second traffic sequence to obtain a second feature tensor; Based on the output layer, feature fusion is performed on the first feature tensor and the second feature tensor, and a traffic prediction value for each base station in a time period to be predicted is output; Among them, the first traffic sequence is the continuous traffic data of multiple time periods and multiple base stations before the time period to be predicted; the second traffic sequence is the continuous traffic data of multiple time periods and the same multiple base stations before and after the same time period on the day before the time period to be predicted.
[0005] Optionally, the first branch includes a one-dimensional convolutional network, a temporal self-attention long short-term memory module, and a spatial multi-head attention module; The first branch extracts the spatial correlation between the traffic of adjacent base stations and the temporal correlation of the traffic sequence of a single base station based on the first traffic sequence, and uses the self-attention mechanism in the time dimension to model the sudden change of the traffic of the same base station in time, and uses the multi-head attention mechanism in the spatial dimension to model the mutation distribution law of the traffic between different base stations in the same period, and obtains the first feature tensor, including: The one-dimensional convolutional network extracts the spatial correlation features between the traffic flows of adjacent base stations based on the first traffic sequence to obtain the spatial correlation features of each time period; The temporal self-attention long short-term memory module extracts the temporal correlation characteristics of the traffic in the previous and next time periods based on the spatial correlation characteristics of each time period to obtain a spatiotemporal feature tensor; The spatial multi-head attention module calculates the first feature tensor based on the spatiotemporal feature tensor and the spatial correlation features.
[0006] Optionally, the temporal self-attention long short-term memory module includes a first long short-term memory network, a second long short-term memory network and a temporal self-attention module; The temporal self-attention long short-term memory module extracts the temporal correlation characteristics of the traffic in the previous and next time periods based on the spatial correlation characteristics of each time period to obtain a spatiotemporal feature tensor, including: The first long short-term memory network determines the spatiotemporal correlation feature of the current period based on the spatial correlation feature of the current period and the spatiotemporal correlation feature of the previous period, and concatenates the spatiotemporal correlation features of each period to obtain a first output; The temporal self-attention module calculates an attention matrix based on the first output; The second long short-term memory network encodes the attention matrix to obtain the spatiotemporal feature tensor.
[0007] Optionally, the spatial multi-head attention module calculates the first feature tensor based on the spatiotemporal feature tensor and the spatial correlation feature, including: Splitting the spatiotemporal feature tensor into a plurality of heads in the direction of the feature dimension; Calculate a query matrix, a key matrix, and a value matrix corresponding to each head according to the spatiotemporal feature tensor and the spatial correlation features; The first feature tensor is determined according to the query matrix, the key matrix and the value matrix of each head.
[0008] Optionally, the spatial multi-head attention module determines the first feature tensor based on the spatiotemporal feature tensor and the spatial correlation feature by the following formula: in, denote the query matrix, key matrix and value matrix respectively, , , Respectively represent The query matrix, key matrix and value matrix of each head, Indicates the number of heads, , , , represents the trainable linear transformation weights, Represents spatial correlation features, represents the spatiotemporal feature tensor, Indicates The attention matrix of each head, represents the attention calculation function, Represents the result of the multi-head attention mechanism, Represents a splicing operation, represents the first eigentensor.
[0009] Optionally, the second branch is a bidirectional long short-term memory network, and the second feature tensor includes a forward time correlation feature and a reverse time correlation feature; The second branch extracts the time correlation of the second traffic sequence to obtain a second feature tensor, including: The time correlation of the second traffic sequence is extracted from both the forward and reverse directions based on the bidirectional long short-term memory network, and a forward time correlation feature and a reverse time correlation feature are output.
[0010] In a second aspect, the present invention further provides a base station traffic prediction device, comprising the following modules: An acquisition module, used to acquire a flow temporal and spatial sequence and a base station flow prediction model; wherein the flow temporal and spatial sequence includes a first flow sequence and a second flow sequence, and the base station flow prediction model includes a first branch, a second branch and an output layer; A modeling module is used for extracting the spatial correlation between the traffic of adjacent base stations and the temporal correlation of the traffic sequence of a single base station based on the first traffic sequence in the first branch, and using the self-attention mechanism in the time dimension to model the sudden change of the traffic of the same base station in time, and using the multi-head attention mechanism in the spatial dimension to model the mutation distribution law of the traffic between different base stations in the same period, so as to obtain a first feature tensor; An extraction module, used for extracting the time correlation of the second flow sequence in a second branch to obtain a second feature tensor; A prediction module is used to perform feature fusion on the first feature tensor and the second feature tensor based on the output layer, and output the traffic prediction value for each base station in the time period to be predicted; wherein the first traffic sequence is continuous traffic data of multiple time periods and multiple base stations before the time period to be predicted; and the second traffic sequence is continuous traffic data of multiple time periods and the same multiple base stations before and after the same time period on the day before the time period to be predicted.
[0011] In a third aspect, the present invention further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the computer program, the base station traffic prediction method as described in the first aspect above is implemented.
[0012] In a fourth aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the base station traffic prediction method as described in the first aspect above.
[0013] In a fifth aspect, the present invention further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the base station traffic prediction method as described in the first aspect above.
[0014] Based on the above technical solution, the beneficial effects of the present invention compared with the prior art are as follows: The base station traffic prediction method and device provided by the present invention, aiming at the spatiotemporal correlation characteristics of the railway 5G private network base station traffic, extracts the spatial correlation between adjacent base station traffic through the first branch, and extracts the temporal correlation of the base station traffic sequence; aiming at the burstiness of the railway 5G private network base station traffic, uses the self-attention mechanism in the time dimension to model the sudden changes in the traffic of the same base station in time, and uses the multi-head attention mechanism in the space dimension to model the mutation distribution law of the traffic between different base stations in the same period; aiming at the periodic characteristics of the railway 5G private network base station traffic, extracts the temporal correlation of the periodic adjacent traffic sequence through the second branch; finally, the feature tensors encoded by the two branches are fused, and the traffic value to be predicted is output by the fully connected layer. By combining spatial correlation, sudden changes in time and periodic changes, the method can capture the changing law of base station traffic more comprehensively. The use of self-attention mechanism and multi-head attention mechanism can more effectively handle the complex relationship in spatiotemporal sequence data and improve the prediction accuracy.
[0015] Other features and advantages of the present invention will be described in the following description, and partly become apparent from the description, or understood by practicing the present invention. The purpose and other advantages of the present invention are realized and obtained by the structures particularly pointed out in the description, claims and drawings.
[0016] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 It is a flow chart of the base station traffic prediction method provided by the present invention.
[0019] Figure 2 It is a structural schematic diagram of the base station traffic prediction model provided by the present invention.
[0020] Figure 3 It is a structural schematic diagram of the temporal self-attention long-term short-term memory module provided by the present invention.
[0021] Figure 4 It is a structural schematic diagram of the temporal self-attention module provided by the present invention.
[0022] Figure 5It is a structural schematic diagram of the spatial multi-head attention module provided by the present invention.
[0023] Figure 6 It is a schematic diagram of the spatial multi-head attention calculation process provided by the present invention.
[0024] Figure 7 It is a structural schematic diagram of a bidirectional long short-term memory network provided by the present invention.
[0025] Figure 8 It is a structural schematic diagram of the base station traffic prediction device provided by the present invention.
[0026] Fig. 9 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0027] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0028] Most of the existing base station traffic prediction technologies are designed for the business scenarios and rules of public cellular base stations. The business negative traffic prediction problems of public base stations are divided into two main types - time series prediction problems and spatiotemporal series prediction problems. For the problem of single base station traffic time series prediction, since there is only one base station, only the traffic of users or devices connected to the base station is considered, and only the time dependency in the historical traffic data of a single base station is used to predict the future traffic of the base station. For the more complex problem of multi-base station traffic spatiotemporal series prediction, the traffic of the base station is not only related to the users or devices connected to the base station, but also to the movement and switching of users between base stations. Therefore, in addition to considering the time correlation of the traffic value sequence of a single base station, it is also necessary to consider the traffic and its spatial correlation in multiple base stations or multiple regions.
[0029] Existing base station traffic prediction models can be roughly divided into three categories, namely statistical models, machine learning models, and deep learning models. Most statistical models are based on the linear relationship between input and output values, and their performance often lags behind that of machine learning models that have the ability to describe and learn nonlinear relationships. Due to the use of manually extracted and constructed features and the limitations of the solution process, machine learning models have limited ability to learn the complex spatiotemporal correlations of traffic data, and often do not work well in base station traffic prediction. At present, most of the traffic prediction methods with better results are deep learning models. Models based on deep learning have the ability to model spatiotemporal correlations and are widely used in spatiotemporal series prediction problems. The core of the spatiotemporal series prediction model based on deep learning is the mining, extraction and fusion of temporal correlation and spatial correlation. The method commonly used for temporal correlation modeling is recurrent neural network. According to the characteristics of the spatial dimension of spatiotemporal sequences, for spatiotemporal sequences with a grid-like spatial structure, convolutional neural networks are used to extract spatial correlations; for spatiotemporal sequences with a graph structure, graph convolutional networks are used to extract spatial correlations.
[0030] The unique spatiotemporal characteristics of the base station traffic of the railway 5G private network have brought new challenges to the establishment of the traffic prediction model. From the spatial dimension, the base stations are arranged in a one-dimensional linear manner along the railway. From the temporal dimension, the changes in the base station traffic have a periodicity of a day and are also sudden. The traffic values in adjacent time periods do not fluctuate smoothly, but show an irregular sawtooth shape with alternating peaks and valleys. This makes the effect of using the conventional public network base station service load prediction method unsatisfactory. The attention mechanism is a method of simulating human attention thinking. By giving the model the ability to focus, the model can improve its learning and utilization of important information. In spatiotemporal sequence prediction, the attention mechanism can help the model better capture the key spatiotemporal features in the sequence and improve the prediction accuracy. By adjusting the attention weights of different parts, the model can dynamically focus on different parts of the input sequence, making the prediction results more accurate and reliable. Based on the above background, the present invention proposes a base station traffic prediction method based on spatiotemporal attention and multi-branch convolution-long short-term memory network in a railway 5G private network.
[0031] The nouns involved in the present invention are explained below: (1) Railway 5G private network: The railway 5G private network mainly carries services related to train command, control, operation and maintenance. The business data types include voice, data, and video. After superimposing its application business scenarios, it includes railway emergency calls, train-related dispatching voice communications, other dispatching voice communications, dispatching video communications, train information data transmission, train safety data transmission, operation and maintenance voice communications, operation and maintenance alarm data transmission, operation and maintenance control data transmission, operation and maintenance other data transmission, operation and maintenance video transmission, emergency voice communications, emergency data transmission, emergency video communications, emergency video transmission, etc.
[0032] (2) Periodicity of traffic (the periodicity of train operation leads to the periodicity of base station traffic): Since the business of base stations in the railway 5G private network mainly comes from the interaction between the base stations and passing trains, the changes in base station traffic are inevitably closely related to the running rules of trains. Because the daily running schedules of trains are often similar, the traffic of base stations in the railway 5G private network also has a daily periodicity.
[0033] (3) Traffic burstiness: In the usual public network scenario, the traffic changes of the same base station are often continuous and gentle fluctuations, and the closer the base stations are to each other at the same time, the closer the traffic values are. However, in the railway 5G private network, since business traffic is generated only when the train passes by the base station and interacts with the base station, the traffic changes of the base station are obviously sudden. From the time dimension, the traffic changes of the same base station are not gentle fluctuations but sudden, showing an irregular sawtooth shape with alternating peaks and valleys; from the spatial dimension, the distribution of traffic of different base stations at the same time is also different from that in the conventional public network scenario. Adjacent base stations are not necessarily close in traffic, and the traffic distribution shows a sudden change of some high and some low.
[0034] Combine the following Figure 1-Figure 8 The base station traffic prediction method and device of the present invention are described.
[0035] Reference Figure 1 As shown, the base station traffic prediction method includes the following: Step S110, obtaining a traffic spatiotemporal sequence and a base station traffic prediction model; wherein the traffic spatiotemporal sequence includes a first traffic sequence and a second traffic sequence, and the base station traffic prediction model includes a first branch, a second branch and an output layer.
[0036] The first traffic sequence is continuous traffic data of multiple time periods and multiple base stations before the time period to be predicted; the second traffic sequence is continuous traffic data of multiple time periods and the same multiple base stations before and after the same time period one day before the time period to be predicted.
[0037] Obtain the traffic spatiotemporal sequence, including the first traffic sequence and the second traffic sequence. Obtain the first traffic sequence, specifically, collect continuous traffic data of multiple time periods and multiple base stations before the time period to be predicted. Continuous time period Traffic of base stations These data reflect the historical changes in base station traffic and are an important basis for predicting future traffic. To obtain the second traffic sequence, specifically, collect continuous traffic data of the same multiple base stations in multiple periods before and after the same period of the day before the predicted period. Continuous time period Traffic of base stations These data help capture the periodic changes in base station traffic and improve the accuracy of predictions.
[0038] Data cleaning is performed on the first and second flow series to remove missing values or outliers. For missing values, interpolation methods (such as linear interpolation, Newton interpolation, etc.) can be used to fill them; for outliers, statistical methods (such as the 3σ principle) or machine learning algorithms (such as isolation forests) can be used to identify and eliminate them. The flow data is converted into a uniform range (such as 0-1) to improve the training efficiency and prediction accuracy of the model. Normalization methods include minimum-maximum normalization, Z-score normalization, etc.
[0039] The above-mentioned base station traffic prediction model consists of the first branch, the second branch and the output layer. The first branch is the main branch, and its input is the traffic flow before the prediction period. Continuous time period Traffic of base stations The first branch is responsible for extracting spatial correlation and temporal sudden changes. The second branch is an auxiliary branch, whose input is the data before and after the same period of the day before the period to be predicted. Continuous time period Traffic of base stations ,The second branch is responsible for extracting temporal correlation, and the output layer is responsible for feature fusion and traffic prediction.
[0040] The preprocessed traffic spatiotemporal sequence is input into the trained base station traffic prediction model, and the model can output the traffic prediction value of each base station in the prediction period according to the input data.
[0041] Step S120: The first branch extracts the spatial correlation between the traffic of adjacent base stations and the temporal correlation of the traffic sequence of a single base station based on the first traffic sequence, and uses the self-attention mechanism in the time dimension to model the sudden change of the traffic of the same base station in time, and uses the multi-head attention mechanism in the spatial dimension to model the mutation distribution law of the traffic between different base stations in the same period, so as to obtain the first feature tensor; The first branch processes the first traffic sequence and extracts the spatial correlation between the traffic of adjacent base stations. Graph neural networks can process data in non-Euclidean space, such as the connection relationship between base stations, thereby capturing the spatial distribution law of base station traffic.
[0042] The first branch also models the sudden changes in time. Specifically, in the time dimension, the self-attention mechanism is used to process the first traffic sequence to model the sudden changes in the traffic of the same base station in time. The self-attention mechanism can dynamically capture the key information in the time series and improve the prediction ability of the model.
[0043] The first branch also models the spatial mutation distribution law. Specifically, in the spatial dimension, a multi-head attention mechanism is used to process the traffic between different base stations in the same period to model the mutation distribution law of traffic. The multi-head attention mechanism can focus on the information of multiple locations at the same time, enhancing the model's ability to handle complex relationships.
[0044] Step S130: The second branch extracts the time correlation of the second flow sequence to obtain a second feature tensor.
[0045] The second branch can use a recurrent neural network (such as LSTM or GRU) to process the second traffic sequence and extract temporal correlation. Recurrent neural networks can capture long-term dependencies in time series and help predict future traffic changes.
[0046] Step S140: Based on the output layer, feature fusion is performed on the first feature tensor and the second feature tensor, and a traffic prediction value for each base station in the time period to be predicted is output.
[0047] The output layer includes a feature fusion layer and a fully connected layer. The feature fusion layer concatenates or weighted fuses the first feature tensor and the second feature tensor to obtain a fused feature vector. Concatenation is to concatenate two feature tensors by dimension, and weighted fusion is to weighted sum the two feature tensors according to the weight. The fused feature vector is processed using a fully connected layer to output the traffic prediction value of each base station during the period to be predicted. The fully connected layer can map the fused feature vector to the prediction value to achieve traffic prediction.
[0048] The base station traffic prediction method provided by the present invention, in view of the spatiotemporal correlation characteristics of the railway 5G private network base station traffic, extracts the spatial correlation between adjacent base station traffic through the first branch, and extracts the temporal correlation of the base station traffic sequence; in view of the burstiness of the railway 5G private network base station traffic, the self-attention mechanism is used in the time dimension to model the sudden changes in the traffic of the same base station in time, and the multi-head attention mechanism is used in the spatial dimension to model the mutation distribution law of the traffic between different base stations in the same period; in view of the periodic characteristics of the railway 5G private network base station traffic, the second branch is used to extract the temporal correlation of the periodic adjacent traffic sequence; finally, the feature tensors encoded by the two branches are fused, and the fully connected layer outputs the traffic value to be predicted. By combining spatial correlation, sudden changes in time and periodic changes, the method can capture the changing law of base station traffic more comprehensively. The use of self-attention mechanism and multi-head attention mechanism can more effectively handle the complex relationships in spatiotemporal sequence data and improve the prediction accuracy.
[0049] Moreover, this method can process traffic data from different base stations and different time periods, and has strong generalization ability. Through the pre-trained traffic prediction model, it can quickly adapt to new base station traffic prediction tasks. Accurate base station traffic prediction helps operators dynamically adjust network resource allocation according to traffic demand. Shutting down some base stations or carrier frequencies during low traffic periods can reduce energy consumption and operating costs. By predicting base station traffic, operators can deploy network resources in advance to ensure stable network services during peak hours and improve user experience and satisfaction.
[0050] In an optional embodiment, the main task of the first branch is to extract the spatial correlation between the traffic of adjacent base stations from the first traffic sequence, and use the self-attention mechanism in the time dimension to model the sudden changes in the traffic of the same base station in time, and use the multi-head attention mechanism in the spatial dimension to model the sudden distribution law of the traffic between different base stations in the same period, and finally obtain the first feature tensor. Figure 2 As shown, the first branch includes a one-dimensional convolutional network (1D-CNN), a temporal self-attention long-term memory module (i.e. Figure 2 The temporal self-attention LSTM network in ), the spatial multi-head attention module.
[0051] The first branch described in step S120 extracts the spatial correlation between the traffic of adjacent base stations and the temporal correlation of the traffic sequence of a single base station based on the first traffic sequence, and uses the self-attention mechanism in the time dimension to model the sudden changes in the traffic of the same base station in time, and uses the multi-head attention mechanism in the spatial dimension to model the mutation distribution law of the traffic between different base stations in the same period, and obtains the first feature tensor, including: S1201. A one-dimensional convolutional network extracts spatial correlation features between traffic flows of adjacent base stations based on the first traffic sequence to obtain spatial correlation features for each time period.
[0052] The one-dimensional convolutional network is used to process the first traffic sequence, that is, the continuous traffic data of multiple time periods and multiple base stations before the predicted period. Through the convolution operation, the network can extract the spatial correlation characteristics between the traffic of adjacent base stations. Specifically, the convolution kernel slides on the traffic data, locally perceives the base station traffic in each time period, and thus extracts the spatial correlation features of each time period. These features reflect the relationship between the traffic of adjacent base stations and provide a basis for subsequent temporal self-attention and spatial multi-head attention processing.
[0053] S1202. The temporal self-attention short-term memory module extracts the temporal correlation characteristics of the traffic in the previous and next time periods based on the spatial correlation characteristics of each time period to obtain a spatiotemporal feature tensor.
[0054] The Temporal Self-Attention LSTM Module is used to extract spatiotemporal features. The Temporal Self-Attention LSTM Module combines the advantages of the self-attention mechanism and the long short-term memory network (LSTM). The module first receives the spatial correlation features of each time period output by the one-dimensional convolutional network, and then uses the self-attention mechanism to capture the temporal correlation features of the traffic in the previous and next time periods. The self-attention mechanism dynamically adjusts the influence weight of each time period on the current time period by calculating the similarity between different time periods, thereby capturing the temporal sudden changes in traffic. At the same time, the LSTM network can capture the long-term dependencies in the time series and further extract the spatiotemporal feature tensors. These feature tensors contain both spatial correlation information and temporal correlation information, providing rich feature representation for subsequent spatial multi-head attention processing.
[0055] S1203. The spatial multi-head attention module calculates the first feature tensor based on the spatiotemporal feature tensor and the spatial correlation features.
[0056] The Spatial Multi-Head Attention Module is based on the spatial correlation features of the spatiotemporal feature tensor and the output of the one-dimensional convolutional network, and calculates the first feature tensor through the spatial multi-head attention mechanism. The spatial multi-head attention mechanism divides the input features into multiple subspaces, and performs attention calculations independently in each subspace, and finally splices the attention results of all subspaces to obtain the final attention output. This mechanism can pay attention to information at multiple locations at the same time, enhancing the model's ability to handle complex relationships. Through the processing of the spatial multi-head attention module, the first feature tensor not only contains the spatial correlation information and time-sudden change information of the base station traffic, but also contains the information on the mutation distribution law of the traffic between different base stations in the same period, providing a comprehensive feature representation for subsequent traffic prediction.
[0057] The present invention combines a one-dimensional convolutional network, a temporal self-attention long and short-term memory module, and a spatial multi-head attention module. The first branch of the present invention can fully capture the spatial correlation of base station traffic, temporal sudden changes, and the mutation distribution law of traffic between different base stations in the same period, thereby improving the accuracy of the prediction. The design of the first branch takes into account the spatiotemporal characteristics of base station traffic, so that the model can process traffic data of different base stations and different time periods, and has a strong generalization ability. This helps the model to quickly adapt to new base station traffic prediction tasks and improves the practicality and flexibility of the model. The spatial correlation features are extracted through a one-dimensional convolutional network, reducing the amount of calculation for subsequent processing; at the same time, the combination of the temporal self-attention long and short-term memory module and the spatial multi-head attention module enables the model to efficiently process large-scale base station traffic data, improving the calculation efficiency.
[0058] In an optional embodiment, the temporal self-attention long short-term memory module in the above step S1202 includes a first long short-term memory network, a second long short-term memory network and a temporal self-attention module; The temporal self-attention long short-term memory module described in the above step S1202 extracts the temporal correlation characteristics of the traffic in the previous and next time periods based on the spatial correlation characteristics of each time period to obtain a spatiotemporal feature tensor, including: S12021. The first long short-term memory network determines the spatiotemporal correlation features of the current time period based on the spatial correlation features of the current time period and the spatiotemporal correlation features of the previous time period, and concatenates the spatiotemporal correlation features of each time period to obtain a first output.
[0059] The first long short-term memory network (LSTM1) receives the spatial correlation features of the current period and the spatiotemporal correlation features of the previous period as input. The spatial correlation features are extracted from the first traffic sequence by a one-dimensional convolutional network, reflecting the relationship between the traffic of adjacent base stations. The spatiotemporal correlation features of the previous period are the output of the module after processing in the previous period, which contains the spatial and temporal correlation information of the previous period.
[0060] LSTM1 fuses and processes the spatial correlation features of the current period and the spatiotemporal correlation features of the previous period through its internal forget gate, input gate, and output gate mechanism to determine the spatiotemporal correlation features of the current period. This step realizes the conversion from spatial correlation features to spatiotemporal correlation features, and takes into account the continuity of the time series. Subsequently, LSTM1 concatenates the spatiotemporal correlation features of each period to obtain the first output. This output contains the spatiotemporal correlation features of all periods, providing a basis for subsequent temporal self-attention processing.
[0061] S12022. The temporal self-attention module calculates an attention matrix based on the first output.
[0062] The Temporal Self-Attention Module receives the first output of the first LSTM network as input. The module captures the temporal correlation characteristics of traffic in previous and subsequent time periods by calculating the similarity between different time periods. Specifically, the Temporal Self-Attention Module first calculates an attention matrix that reflects the degree of correlation between different time periods. Then, the module uses this attention matrix to weight the first output to obtain a weighted feature representation. This step realizes the dynamic capture and emphasis of key information in the time series.
[0063] S12023. The second long short-term memory network encodes the attention matrix to obtain the spatiotemporal feature tensor.
[0064] The structure of the second LSTM network is the same as that of the first LSTM network. The second LSTM network (LSTM2) receives the weighted feature representation output by the temporal self-attention module as input. LSTM2 further encodes and processes the weighted feature representation through its internal gating mechanism to extract deeper spatiotemporal features. The output of LSTM2 is the final spatiotemporal feature tensor, which contains both spatial correlation information and temporal correlation information, providing a comprehensive feature representation for subsequent traffic prediction.
[0065] By combining the long short-term memory network (LSTM) and the temporal self-attention mechanism, the module can fully capture the spatiotemporal characteristics of base station traffic, including spatial correlation, temporal continuity, and temporal burst changes. This enables the model to more accurately capture key information when predicting base station traffic, improving the accuracy and reliability of the prediction. The design of the temporal self-attention long short-term memory module takes into account the spatiotemporal characteristics of base station traffic, enabling the model to process traffic data from different base stations and different time periods. This helps the model quickly adapt to new base station traffic prediction tasks and improves the practicality and flexibility of the model. Moreover, through the combination of the gating mechanism of the long short-term memory network and the temporal self-attention mechanism, the module can efficiently process large-scale base station traffic data. This reduces the consumption of computing resources, improves computing efficiency, and enables the model to complete the prediction task in a shorter time.
[0066] In an optional embodiment, the spatial multi-head attention module described in step S1203 calculates the first feature tensor based on the spatiotemporal feature tensor and the spatial correlation feature, including: S12031. Divide the spatiotemporal feature tensor into multiple heads in the direction of the feature dimension.
[0067] First, the spatial multi-head attention module receives the spatiotemporal feature tensor and spatial correlation features as input. In order to process and capture information in different subspaces in parallel, the module splits the spatiotemporal feature tensor into multiple heads in the direction of the feature dimension (such as the last dimension of the tensor). Each head is a smaller subset of features, which will be independently used for subsequent attention calculations.
[0068] S12032. Calculate the query matrix, key matrix and value matrix corresponding to each head according to the spatiotemporal feature tensor and the spatial correlation features.
[0069] For each head, the spatial multi-head attention module calculates its corresponding query matrix (Query), key matrix (Key), and value matrix (Value). These matrices are obtained by applying linear transformations to the spatiotemporal feature tensors and spatially related features. Specifically, for each head, there are three different linear transformation matrices, which are used to generate the query matrix, key matrix, and value matrix respectively.
[0070] S12033. Determine the first feature tensor according to the query matrix, key matrix, and value matrix of each head.
[0071] After obtaining the query matrix, key matrix, and value matrix for each head, the spatial multi-head attention module performs standard attention calculations. This involves calculating the dot product of the query matrix and the key matrix, then applying the softmax function to obtain the attention weights, and finally using these weights to perform a weighted sum of the value matrix. This step generates an output for each head .
[0072] The attention matrices of all heads are concatenated in the direction of the feature dimension and integrated through an additional linear transformation to finally obtain the first feature tensor. This tensor contains the attention features extracted from multiple subspaces, providing rich information for subsequent traffic prediction.
[0073] The present invention divides the features into multiple heads and performs attention calculations independently on each head, so that the module can capture information in different subspaces. This enables the model to focus on different positions in the sequence at the same time and extract diverse features, thereby enhancing the representation ability of the model. The multi-head attention mechanism allows the model to process information in multiple subspaces in parallel, thereby more comprehensively capturing the spatiotemporal characteristics of base station traffic. This helps the model to more accurately capture key information when predicting base station traffic and improve the accuracy and reliability of the prediction. Since each head performs attention calculations independently, they can capture different features and information. This makes the model less sensitive to slight changes in the input data, thereby enhancing the robustness and stability of the model. The multi-head attention mechanism allows the model to process multiple heads in parallel, thereby improving computational efficiency. This enables the model to complete the prediction task in a shorter time and reduces the consumption of computing resources.
[0074] In an optional embodiment, the second branch described in step S130 is a bidirectional long short-term memory network, and the second feature tensor includes a forward time correlation feature and a reverse time correlation feature; The second branch described in step S130 extracts the time correlation of the second flow sequence to obtain a second feature tensor, including: S1301. Extract the time correlation of the second traffic sequence from both the forward and reverse directions based on the bidirectional long short-term memory network, and output forward time correlation features and reverse time correlation features.
[0075] The second traffic sequence is used as the input of the bidirectional long short-term memory network. This sequence represents a series of base station traffic data arranged in chronological order. Figure 7 As shown in Figure 1, the bidirectional long short-term memory network consists of two independent LSTM networks, one processing the sequence from front to back (forward LSTM) and the other processing the sequence from back to front (reverse LSTM). The two LSTM networks share the same input sequence but process in opposite directions.
[0076] The forward LSTM calculates sequentially from the first element to the last element of the sequence and extracts the forward time correlation features. These features reflect the time dependency from the beginning to the end of the sequence. The reverse LSTM calculates in reverse order from the last element to the first element of the sequence and extracts the reverse time correlation features. These features reflect the time dependency from the end to the beginning of the sequence. The outputs of the forward LSTM and the reverse LSTM are concatenated to form the final output tensor. This tensor contains both the forward time correlation features and the reverse time correlation features.
[0077] The present invention adopts a bidirectional long short-term memory network to simultaneously capture the time dependencies from front to back and from back to front in the sequence. This enables the model to more comprehensively consider the contextual information of the time series when predicting base station traffic, thereby improving the accuracy and reliability of the prediction. By combining the LSTM outputs in both the forward and reverse directions, the model can extract richer feature information. This helps the model to better represent the complexity and diversity of time series data, thereby improving the representation ability of the model. The bidirectional long short-term memory network has a certain robustness to slight changes in the input data. Since the model considers both the forward and reverse time dependencies, the model can maintain stable prediction performance to a certain extent even if there is noise or missing values in the input data.
[0078] Specifically, the process of performing traffic prediction based on the base station traffic prediction model of the present invention is as follows: S210, one-dimensional convolutional network extracts the spatial correlation characteristics between adjacent base station traffic, and for each time period Traffic of each base station Perform a one-dimensional convolution operation: (1-1) in, for The output of the one-dimensional convolutional network in the time period, that is, the above-mentioned spatial correlation feature, is the activation function, is the weight of the trainable filter, for Traffic of each base station during the period, symbol represents a one-dimensional convolution operation, is a trainable bias.
[0079] S220, the temporal self-attention long short-term memory module consists of two long short-term memory networks (LTSM networks) and a temporal self-attention module, and its structure is as follows Figure 3 For the first layer of LSTM network (i.e. the first long short-term memory network mentioned above), the time correlation characteristics of the traffic in the previous and next time periods are extracted through the following process: (1-2) ( + )(1-3) (1-4) (1-5) (1-6) in, It is Input of time steps, here input Corresponding to the spatial correlation features output by the previous one-dimensional convolutional network . The above After processing (such as directly passing it or after simple preprocessing), it is used as the input of the first long short-term memory network. and is the activation function, , , They are input gate, forget gate and output gate respectively. yes The spatiotemporal correlation characteristics of the time period, yes The unit status of the time period, yes The unit status of the time period, represents matrix element multiplication, (i.e. the above , , , , , , , , , , )and (i.e. the above , , , ) are trainable weights and biases, is the hyperbolic tangent function.
[0080] Finally, each Spatial-temporal correlation characteristics of time periods Concatenate to get the output of the first layer of LSTM network , which is the first output mentioned above.
[0081] S230. The temporal self-attention module uses the self-attention mechanism in the temporal dimension to model the sudden changes in the traffic of the same base station over time. According to the importance of each time step in the traffic sequence, the traffic at different time steps is weighted. In this way, the model can pay more attention to those time steps of traffic mutations that have a greater impact on the prediction results. The structure of the temporal self-attention module is as shown in Figure 4 Figure. In the temporal self-attention module, the operations MatMul, Scale, Mask, SoftMax, and MatMul from bottom to top respectively represent matrix multiplication, scaling, masking, activation function, and matrix multiplication operations. These operations together implement the temporal self-attention mechanism for modeling the sudden changes in the traffic of the same base station over time.
[0082] In the temporal self-attention mechanism, first, the query matrix (Q), key matrix (K), and value matrix (V) are calculated through matrix multiplication (MatMul). These matrices are obtained by multiplying the output of the first LSTM network with the trainable linear transformation weights , , .
[0083] When calculating dot-product attention, to avoid the problem of gradient vanishing or explosion caused by a large key vector length, a scaling factor is introduced. This scaling operation (Scale) divides the dot-product result of the query matrix Q and the key matrix K by to ensure the stability of training.
[0084] The Mask layer is used to restrict the attention weights. For example, when processing sequence data, it may be desired that the model only pays attention to the elements before the current position and ignores the elements after. This can be achieved by applying a mask on the attention weights.
[0085] The SoftMax function is used to calculate the attention weights to ensure that the sum of the weights is 1. It converts the scaled dot-product result into a probability distribution, thus obtaining the attention weights for each time step.
[0086] Finally, through matrix multiplication (MatMul), the attention weight matrix is multiplied by the value matrix V to obtain the weighted output. This output is the final result of the temporal self-attention module, which weights the traffic according to the importance of each time step.
[0087] The attention matrix is calculated as follows: (1 - 7) (1-8) (1-9) (1-10) in, , , denote the query matrix, key matrix and value matrix respectively, is the output of the first LSTM network, that is, the first output mentioned above. , , is the trainable linear transformation weight, For the dimension of the key vector, introduce a scaling factor This is to avoid the gradient vanishing or exploding problem caused by the large length of the key vector when calculating the dot product, thereby ensuring the stability of training. represents the attention calculation function, It is a function in the attention mechanism that is used to calculate the attention weights and ensure that the sum of the weights is 1.
[0088] The structure of the second LSTM network is the same as the first LSTM network, and the attention matrix The second LSTM network encodes and outputs the spatiotemporal feature tensor .
[0089] S240, the spatial multi-head attention module uses a multi-head attention mechanism in the spatial dimension to model the mutation distribution law of traffic between different base stations in the same period. The structure of the spatial multi-head attention module is as follows Figure 5 shown.
[0090] Reference Figure 5 As shown in the figure, the spatial multi-head attention module includes Linear, Scaled Dot-ProductAttention, Concat, and Linear from bottom to top. Linear refers to different linear transformations. , , When the spatial correlation characteristics or spatiotemporal feature tensor Each head will be transformed through h different linear transformations to obtain , , These linear transformations are represented by trainable linear transformation weights , , definition.
[0091] For each head, the Scaled Dot-Product Attention mechanism is used to calculate the attention matrix . This involves dividing the dot product result of and by ( is the dimension of the key vector), then calculating the attention weights through the softmax function, and finally applying these weights to to obtain the weighted output, which is the above-mentioned attention matrix .
[0092] The attention matrices of all heads are concatenated along the feature dimension to form the result of the multi-head attention mechanism MultiHeadAttention(K,Q,V).
[0093] Finally, the result of the multi-head attention mechanism passes through another linear transformation Linear (defined by the trainable linear transformation weights ) to obtain the final first feature tensor .
[0094] Specifically, the input feature tensor (i.e., the spatio-temporal feature tensor output by the above second LSTM network ) is sliced into h heads in the direction of the feature dimension, and the attention matrix is calculated according to the following formula, which is the above-mentioned first feature tensor : (1-11) (1-12) (1-13) (1-14) (1-15) (1-16) Among them, represent the query matrix, key matrix, and value matrix respectively, , , represent the query matrix, key matrix, and value matrix of the th head respectively, represents the number of heads, , , , represent the trainable linear transformation weights, represents the space-related features, which are the output of the above one-dimensional convolutional network, represents the spatiotemporal feature tensor, Indicates The attention matrix of each head, represents the attention calculation function, It is a function in the attention mechanism that is used to calculate the attention weights and ensure that the sum of the weights is 1. Represents the result of the multi-head attention mechanism, Represents the splicing operation, splicing along the feature dimension direction, represents the first eigentensor.
[0095] Taking the 8-head spatial attention mechanism as an example, Figure 6 The above calculation process is demonstrated.
[0096] S240, such as Figure 7 As shown in Figure 1, the bidirectional LSTM network consists of the same LSTM network in two directions, extracting the second flow sequence from both the forward and reverse directions. Time correlation, output positive time correlation feature and reverse time correlation characteristics .
[0097] S250, in the output layer, the feature tensors output by the first branch and the second branch are first concatenated along the feature dimension direction to perform feature fusion, and then the predicted value of the traffic of each base station in the future period is output through the fully connected layer: (1-17) (1-18) in, represents the fused features, Represents a splicing operation, represents the first eigentensor, represents the positive time correlation feature, represents the reverse time correlation feature, represents the linear transformation weight of the fully connected layer, Represents the output of the output layer, that is, the traffic forecast value for the above-mentioned period to be predicted.
[0098] The present invention is based on the time period to be predicted Continuous time period Traffic of base stations The same period before and after the day before the forecast period Continuous time period Traffic of base stations Predict the traffic of each base station in a future period (i.e. the traffic forecast value of each base station in the predicted period) : (2-1) in, represents the base station traffic prediction model of the present invention, For all parameters to be trained, select the mean square error As a loss function, minimized during model training : (2-2) in, is the traffic forecast value of each base station during the forecast period, is the actual value of the traffic flow of each base station during the period to be predicted, Indicates the number of base stations.
[0099] The base station traffic prediction method proposed in the present invention is a base station traffic prediction method based on spatiotemporal attention and multi-branch convolution-long short-term memory network in railway 5G private network. In view of the spatiotemporal correlation characteristics of the base station traffic of railway 5G private network, a one-dimensional convolutional network is used to extract the spatial correlation between adjacent base station traffic, and a long short-term memory network is used to extract the temporal correlation of the base station traffic sequence; in view of the burstiness of the base station traffic of railway 5G private network, a self-attention mechanism is used in the time dimension to model the sudden changes in the traffic of the same base station in time, and a multi-head attention mechanism is used in the space dimension to model the mutation distribution law of the traffic between different base stations in the same period; in view of the periodic characteristics of the base station traffic of railway 5G private network, a structure with multiple branch inputs is used to process two adjacent historical sequences of cycles to model the periodicity of the traffic. In the second branch, a bidirectional long short-term memory network is used to extract the temporal correlation of the adjacent traffic sequences of the cycle from both the forward and reverse directions, and the periodic features encoded by the second branch are combined with the spatiotemporal correlation features encoded by the first branch through feature fusion, and the predicted value of the traffic of each base station in the period to be predicted is output by the fully connected layer.
[0100] The base station traffic prediction device provided by the present invention is described below. The base station traffic prediction device described below and the base station traffic prediction method described above can be referenced to each other.
[0101] The base station traffic prediction device provided by the present invention refers to Figure 8 As shown, including: The acquisition module 310 is used to acquire a flow temporal and spatial sequence and a base station flow prediction model; wherein the flow temporal and spatial sequence includes a first flow sequence and a second flow sequence, and the base station flow prediction model includes a first branch, a second branch and an output layer; Modeling module 320, for extracting spatial correlation between traffic flows of adjacent base stations based on the first traffic sequence in the first branch, and using a self-attention mechanism in the time dimension to model the sudden change of traffic flows of the same base station in time, and using a multi-head attention mechanism in the spatial dimension to model the sudden distribution law of traffic flows between different base stations in the same period, to obtain a first feature tensor; An extraction module 330, configured to extract the time correlation of the second flow sequence in a second branch to obtain a second feature tensor; The prediction module 340 is used to perform feature fusion on the first feature tensor and the second feature tensor based on the output layer, and output the traffic prediction value for each base station in the time period to be predicted; wherein the first traffic sequence is the continuous traffic data of multiple time periods and multiple base stations before the time period to be predicted; and the second traffic sequence is the continuous traffic data of multiple time periods and the same multiple base stations before and after the same time period on the day before the time period to be predicted.
[0102] Fig. 9 An example of a physical structure diagram of an electronic device is shown in FIG. Fig. 9 As shown, the electronic device may include: a processor (processor) 410, a communication interface (Communications Interface) 420, a memory (memory) 430 and a communication bus 440, wherein the processor 410, the communication interface 420, and the memory 430 communicate with each other through the communication bus 440. The processor 410 may call the logic instructions in the memory 430 to execute the base station traffic prediction method.
[0103] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0104] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the base station traffic prediction method provided by the above-mentioned methods.
[0105] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the computer program is executed by a processor to execute the base station traffic prediction method provided by the above-mentioned methods.
[0106] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0107] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0108] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A base station traffic prediction method, characterized in that: include: Acquire a traffic spatiotemporal sequence and a base station traffic prediction model; wherein the traffic spatiotemporal sequence includes a first traffic sequence and a second traffic sequence, and the base station traffic prediction model includes a first branch, a second branch, and an output layer; The first branch extracts the spatial correlation between the traffic of adjacent base stations and the temporal correlation of the traffic sequence of a single base station based on the first traffic sequence, and uses the self-attention mechanism in the time dimension to model the sudden change of the traffic of the same base station in time, and uses the multi-head attention mechanism in the spatial dimension to model the mutation distribution law of the traffic between different base stations in the same period, so as to obtain the first feature tensor; The second branch extracts the time correlation of the second flow sequence to obtain a second feature tensor; Based on the output layer, feature fusion is performed on the first feature tensor and the second feature tensor, and a traffic prediction value for each base station in a time period to be predicted is output; Among them, the first traffic sequence is the continuous traffic data of multiple time periods and multiple base stations before the time period to be predicted; the second traffic sequence is the continuous traffic data of multiple time periods and the same multiple base stations before and after the same time period on the day before the time period to be predicted.
2. The base station traffic prediction method according to claim 1, characterized in that: The first branch includes a one-dimensional convolutional network, a temporal self-attention long and short-term memory module, and a spatial multi-head attention module; The first branch extracts the spatial correlation between the traffic of adjacent base stations and the temporal correlation of the traffic sequence of a single base station based on the first traffic sequence, and uses the self-attention mechanism in the time dimension to model the sudden change of the traffic of the same base station in time, and uses the multi-head attention mechanism in the spatial dimension to model the mutation distribution law of the traffic between different base stations in the same period, and obtains the first feature tensor, including: The one-dimensional convolutional network extracts the spatial correlation features between the traffic flows of adjacent base stations based on the first traffic sequence to obtain the spatial correlation features of each time period; The temporal self-attention long short-term memory module extracts the temporal correlation characteristics of the traffic in the previous and next time periods based on the spatial correlation characteristics of each time period to obtain a spatiotemporal feature tensor; The spatial multi-head attention module calculates the first feature tensor based on the spatiotemporal feature tensor and the spatial correlation features.
3. The base station traffic prediction method according to claim 2, characterized in that: The temporal self-attention long and short-term memory module includes a first long and short-term memory network, a second long and short-term memory network and a temporal self-attention module; The temporal self-attention long short-term memory module extracts the temporal correlation characteristics of the traffic in the previous and next time periods based on the spatial correlation characteristics of each time period to obtain a spatiotemporal feature tensor, including: The first long short-term memory network determines the spatiotemporal correlation feature of the current period based on the spatial correlation feature of the current period and the spatiotemporal correlation feature of the previous period, and concatenates the spatiotemporal correlation features of each period to obtain a first output; The temporal self-attention module calculates an attention matrix based on the first output; The second long short-term memory network encodes the attention matrix to obtain the spatiotemporal feature tensor.
4. The base station traffic prediction method according to claim 2, characterized in that: The spatial multi-head attention module calculates the first feature tensor based on the spatiotemporal feature tensor and the spatial correlation feature, including: Splitting the spatiotemporal feature tensor into a plurality of heads in the direction of the feature dimension; Calculate a query matrix, a key matrix, and a value matrix corresponding to each head according to the spatiotemporal feature tensor and the spatial correlation features; The first feature tensor is determined according to the query matrix, the key matrix and the value matrix of each head.
5. The base station traffic prediction method according to claim 4, characterized in that: The spatial multi-head attention module determines the first feature tensor based on the spatiotemporal feature tensor and the spatial correlation feature by the following formula: in, denote the query matrix, key matrix and value matrix respectively, , , Respectively represent The query matrix, key matrix and value matrix of each head, Indicates the number of heads, , , , represents the trainable linear transformation weights, Represents spatial correlation features, represents the spatiotemporal feature tensor, Indicates The attention matrix of each head, represents the attention calculation function, Represents the result of the multi-head attention mechanism, Represents a splicing operation, represents the first eigentensor.
6. The base station traffic prediction method according to claim 1, characterized in that: The second branch is a bidirectional long short-term memory network, and the second feature tensor includes a forward time correlation feature and a reverse time correlation feature; The second branch extracts the time correlation of the second traffic sequence to obtain a second feature tensor, including: The time correlation of the second traffic sequence is extracted from both the forward and reverse directions based on the bidirectional long short-term memory network, and a forward time correlation feature and a reverse time correlation feature are output.
7. A base station traffic prediction device, characterized in that: include: An acquisition module, used to acquire a flow temporal and spatial sequence and a base station flow prediction model; wherein the flow temporal and spatial sequence includes a first flow sequence and a second flow sequence, and the base station flow prediction model includes a first branch, a second branch and an output layer; A modeling module is used for extracting the spatial correlation between the traffic of adjacent base stations and the temporal correlation of the traffic sequence of a single base station based on the first traffic sequence in the first branch, and using the self-attention mechanism in the time dimension to model the sudden change of the traffic of the same base station in time, and using the multi-head attention mechanism in the spatial dimension to model the mutation distribution law of the traffic between different base stations in the same period, so as to obtain a first feature tensor; An extraction module, used for extracting the time correlation of the second flow sequence in a second branch to obtain a second feature tensor; A prediction module is used to perform feature fusion on the first feature tensor and the second feature tensor based on the output layer, and output the traffic prediction value for each base station in the time period to be predicted; wherein the first traffic sequence is continuous traffic data of multiple time periods and multiple base stations before the time period to be predicted; and the second traffic sequence is continuous traffic data of multiple time periods and the same multiple base stations before and after the same time period on the day before the time period to be predicted.
8. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the base station traffic prediction method according to any one of claims 1 to 6 is implemented.
9. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the base station traffic prediction method according to any one of claims 1 to 6 is implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the base station traffic prediction method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Method, system and device for predicting cellular flow and medium
CN114039871A
Base station flow prediction method and system based on double attention mechanism
CN115589369A
Output traffic prediction method and device, and storage medium
CN116827809A
Air route network flow prediction method based on spatio-temporal feature fusion
CN117058927A
Method and system for traffic prediction based on space-time relation
US20110161261A1