Mobile traffic generation method and device oriented to geographic characteristics
Through a geographic feature-driven modeling approach, utilizing a cross-attention module and a denoising network, we address the data bias and insufficient environment modeling issues in existing mobile traffic generation methods, achieving more accurate mobile traffic generation and higher model adaptability.
Patent Information
- Application Number
- CN202510989826.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-10-14
AI Technical Summary
Existing mobile traffic generation methods rely on VAE and GAN architectures, which cannot fully preserve the key features of the input data and lack explicit environment modeling, resulting in the generated data deviating from the real data and failing to capture the diversity of mobile network traffic patterns.
A geographic feature-driven modeling approach is adopted. By obtaining the point of interest vector and surface of interest vector of the target base station, and using the cross-attention module and denoising network, the dependency between the geographical environment and mobile traffic is adaptively captured to generate a mobile traffic sequence.
It improves the stability and accuracy of the generated model, significantly enhances the modeling ability of complex network traffic data, and improves the adaptability and generalization ability of the model in different regions.
Smart Images

Figure CN120786384A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of mobile communication, in particular to a mobile traffic generation method and device oriented to geographical characteristics. BACKGROUND
[0002] With the rapid development of network communication technology and the popularity of intelligent terminal devices, data traffic has grown explosively, and the resulting base station network load problem has become increasingly prominent. The network load capacity of the base station is closely related to network quality, user experience, etc., and accurate prediction of base station traffic plays an important role in improving network resource utilization and accelerating network intelligent construction.
[0003] The mobile traffic generation method in the related art mainly relies on the variational autoencoder (VAE) and the generative adversarial network (GAN) architecture. The VAE network compresses data into a latent space distribution, but the latent representation may not fully preserve the key features of the input data, especially when dealing with complex traffic data. This may cause the generated data to deviate from the real data in some cases. The GAN-based model requires adversarial training between the generator and the discriminator, which may suffer from mode collapse, i.e., the model generates repetitive patterns and fails to capture the diversity of mobile network traffic patterns.
[0004] Therefore, there is an urgent need for a mobile traffic generation method that can more comprehensively understand the complex dependence between the environment and network traffic. SUMMARY
[0005] The purpose of the present application is to provide a mobile traffic generation method and device oriented to geographical characteristics, which can more accurately model complex network traffic data by using a conditional generation mechanism to adaptively capture the dependence between geographical environment and mobile traffic through a geographical feature-driven modeling approach.
[0006] The present application provides a mobile traffic generation method oriented to geographical characteristics, comprising: obtaining a point of interest vector and an interest surface vector within the signal coverage range of a target base station, extracting traffic sequence features from a randomly generated time domain traffic sequence, and extracting regional environment features from the point of interest vector and the interest surface vector; inputting the regional environment features as keys and values and the traffic sequence features as queries into a cross-attention module to calculate corresponding to-be-processed sequence features, and performing multiple denoising operations on the to-be-processed sequence features to generate a mobile traffic sequence; wherein the point of interest vector is used to represent the number of different types of points of interest within the signal coverage range of the base station; the interest surface vector is used to represent whether each interest surface in a plurality of pre-divided interest surfaces within the range of the base station exists; and the traffic sequence features are used to represent the traffic features of different frequency components.
[0007] Optionally, the target base station is any one of the following: a real base station at a target location, a virtual base station at a target location.
[0008] Optionally, the traffic sequence feature includes: a plurality of frequency domain bases and a time domain residual sequence; and the extracting the traffic sequence feature from the randomly generated time domain traffic sequence includes: converting the time domain traffic sequence from a time domain to a frequency domain by using a discrete Fourier transform to obtain a frequency domain traffic sequence; dividing the frequency domain traffic sequence into a plurality of frequency domain bases and a time domain residual sequence other than the plurality of frequency domain bases; and wherein the plurality of frequency domain bases are used to represent traffic characteristics of different frequency components, and one frequency domain base represents traffic characteristics of one frequency component.
[0009] Optionally, the extracting the regional environment feature from the interest point vector and the interest face vector includes: performing feature extraction on the interest point vector and the interest face vector by using a pre-trained traffic mean prediction period, respectively, and fusing the extracted features based on an attention mechanism to obtain a traffic prediction mean; inputting the traffic prediction mean into a first multi-layer perception to obtain a corresponding traffic mean feature, and performing feature fusion on the traffic mean feature, a location encoding at a current time step, and a result feature obtained by predicting the interest point vector based on a second multi-layer perception to obtain the regional environment feature; and wherein the location encoding is a sine position encoding generated based on a point on the time domain traffic sequence.
[0010] Optionally, the inputting the regional environment feature as a key and a value, and the traffic sequence feature as a query into a cross-attention module to calculate a corresponding to-be-processed feature sequence includes: performing cross-attention calculation on each frequency domain base in the plurality of frequency domain bases as a query with the regional environment feature as a key and a value in the cross-attention module to obtain a first feature corresponding to each frequency domain base; multiplying each frequency domain base corresponding first feature with a corresponding parameter vector to obtain a second feature corresponding to each frequency domain base, and adding the second feature corresponding to each frequency domain base after fusing the second feature to obtain the to-be-processed feature sequence.
[0011] Optionally, the performing multiple denoising operations on the to-be-processed sequence feature to generate a mobile traffic sequence includes: performing multiple denoising operations on the to-be-processed sequence feature by using a residual connection layer in a denoising network to obtain the mobile traffic sequence.
[0012] The application also provides a mobile traffic generation device oriented to geographical characteristics, comprising: The acquisition module is configured to acquire a point-of-interest vector and a surface-of-interest vector in a signal coverage range of a target base station; the feature extraction module is configured to extract a traffic sequence feature from a randomly generated time-domain traffic sequence and extract a regional environment feature from the point-of-interest vector and the surface-of-interest vector; and the traffic generation module is configured to input the regional environment feature as a key and a value and the traffic sequence feature as a query into a cross-attention module, calculate a corresponding to-be-processed sequence feature, and perform multiple denoising operations on the to-be-processed sequence feature to generate a mobile traffic sequence. The point-of-interest vector is used to represent the number of different types of points of interest in the signal coverage range of the base station. The surface-of-interest vector is used to represent whether each surface of interest in a plurality of pre-divided surfaces of interest exists in the range of the base station. The traffic sequence feature is used to represent traffic features of different frequency components.
[0013] Optionally, the traffic sequence feature includes a plurality of frequency-domain bases and a time-domain residual sequence. The feature extraction module is specifically configured to convert the time-domain traffic sequence to a frequency domain by using a discrete Fourier transform to obtain a frequency-domain traffic sequence. The feature extraction module is specifically further configured to divide the frequency-domain traffic sequence into a plurality of frequency-domain bases and a time-domain residual sequence other than the plurality of frequency-domain bases. The plurality of frequency-domain bases are used to represent traffic features of different frequency components, and one frequency-domain base represents traffic features of one frequency component.
[0014] Optionally, the feature extraction module is specifically configured to perform feature extraction on the point-of-interest vector and the surface-of-interest vector by using a pre-trained traffic mean prediction period, and fuse the extracted features based on an attention mechanism to obtain a traffic prediction mean. The feature extraction module is specifically further configured to input the traffic prediction mean into a first multi-layer perception machine to obtain a traffic mean feature, and perform feature fusion on the traffic mean feature, a position encoding of a current time step, and a result feature obtained by predicting the point-of-interest vector based on a second multi-layer perception machine to obtain the regional environment feature. The position encoding is a sine position encoding generated based on a point on the time-domain traffic sequence.
[0015] Optionally, the traffic generation module is specifically configured to perform cross-attention calculation on the regional environment feature as a key and a value and each frequency-domain base in the plurality of frequency-domain bases as a query in a cross-attention module to obtain a first feature corresponding to each frequency-domain base. The traffic generation module is specifically configured to multiply each frequency-domain base corresponding first feature with a corresponding parameter vector to obtain a second feature corresponding to each frequency-domain base, and fuse each frequency-domain base corresponding second feature and add the time-domain residual sequence to obtain the to-be-processed feature sequence.
[0016] Optionally, the traffic generation module is specifically configured to perform multiple denoising operations on the to-be-processed sequence feature by using a residual connection layer in the denoising network to obtain the mobile traffic sequence.
[0017] The application further provides a computer program product, comprising computer programs / instructions, which, when executed by a processor, implement the steps of the mobile traffic generation method according to any one of the above.
[0018] The application further provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the mobile traffic generation method according to any one of the above when executing the program.
[0019] The application further provides a computer-readable storage medium, which stores a computer program, and the computer program, when executed by a processor, implements the steps of the mobile traffic generation method according to any one of the above.
[0020] The application provides a mobile traffic generation method and device oriented to geographical characteristics. First, a point-of-interest vector and an interest area vector in a signal coverage range of a target base station are obtained, and traffic sequence features are extracted from a randomly generated time-domain traffic sequence, and regional environment features are extracted from the point-of-interest vector and the interest area vector. Then, the regional environment features are taken as keys and values, and the traffic sequence features are taken as queries, and input into a cross-attention module, to calculate corresponding to-be-processed sequence features, and perform multiple denoising operations on the to-be-processed sequence features to generate a mobile traffic sequence. The point-of-interest vector is used to represent the number of different types of points of interest in the signal coverage range of the base station. The interest area vector is used to represent whether each interest area in a plurality of pre-divided interest areas in the signal coverage range of the base station exists. The traffic sequence features are used to represent traffic features of different frequency components. In this way, by using a geographical feature driven modeling method and a conditional generation mechanism, the dependence between geographical environment and mobile traffic can be adaptively captured, and complex network traffic data can be more accurately modeled. BRIEF DESCRIPTION OF DRAWINGS
[0021] In order to more clearly illustrate the technical solutions of the present application or the prior art, the following will briefly introduce the drawings needed in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the present application, and those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0022] Figure 1 is a structural schematic diagram of the diffusion model oriented to geographical characteristics provided by the present application. Figure 2 is a flowchart of a mobile traffic generation method for geographical characteristics provided by the present application; Figure 3 is a comparison result diagram of the mobile traffic generation method for geographical characteristics provided by the present application and other algorithms; Figure 4 is a model transferability effect comparison diagram provided by the present application; Figure 5 is a structural diagram of a mobile traffic generation device for geographical characteristics provided by the present application; Figure 6 is a structural diagram of an electronic device provided by the present application. DETAILED DESCRIPTION
[0023] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described clearly and completely below with reference to the drawings in the present application. Obviously, the described embodiments are some embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0024] The terms "first", "second", and the like in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually a category and do not limit the number of objects, for example, the first object can be one or more. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / ", generally means that the front and rear associated objects are in an "or" relationship.
[0025] The mobile network traffic generation technology in the related art faces two main limitations: 1. The existing generation mode has limited capabilities.
[0026] The mobile traffic generation method in the related art mainly relies on VAE and GAN architectures. The VAE network compresses data into a latent space distribution, but the latent representation may not fully preserve the key features of the input data, especially when dealing with complex traffic data. This can cause the generated data to deviate from the real data in some cases. GAN-based models require adversarial training between the generator and the discriminator, which can suffer from mode collapse, i.e., the model generates repetitive patterns and fails to capture the diversity of mobile network traffic patterns.
[0027] 2. Lack of explicit environmental modeling.
[0028] While cities differ in size and population, there are universal patterns that existing models fail to link to the spatiotemporal variations of mobile network traffic. For example, commercial areas typically have more network traffic than parks, and dining areas reach peak usage during lunch and dinner hours. However, existing methods usually map various types of environmental data directly into latent embeddings and rely on the model to implicitly learn the complex relationships between these features. This approach lacks explicit guidance, making it difficult for the model to understand the spatiotemporal associations between mobile network traffic and environmental factors.
[0029] To address the above technical problems existing in the related art, the embodiments of the present application propose a diffusion model (Re-Diff) oriented to geographical characteristics. First, the diffusion model trains a denoising network using a Markov chain, which simulates the noise injection and denoising process. Compared with VAE and GAN architectures, the diffusion model provides higher training stability, effectively alleviating the first limitation. Second, the denoising network is designed to guide the model to learn the dependency between contextual features and traffic patterns. Specifically, a traffic mean predictor is introduced in the denoising process, which is based on the regional point of interest (POI) and area of interest (AOI) distribution, explicitly representing the relationship between the environment and network traffic usage levels. In addition, a cross-attention mechanism is designed to explore the influence of multi-scale temporal patterns on predicted average traffic, enabling the model to understand the temporal dynamics in different spatial contexts, thus addressing the second limitation.
[0030] Focusing on a wireless access network (RAN) composed of multiple base stations (BS) and mobile devices, where a discrete-time system T = {0, 1,..., t} t ∈ [0, T] is considered, with equal time intervals. For a given base station, mobile traffic xt represents the sum of uplink / downlink traffic generated by mobile devices within the coverage of the base station at time t. Meanwhile, the area of interest (AOI) describes the broader functional area where the base station BS is located, such as a commercial area, residential area, industrial area, or office area; the point of interest (POI) distribution P around the base station includes the types and quantities of POIs such as restaurants, shopping centers, transportation facilities, schools, and hospitals, reflecting the functional composition and potential human activity patterns of the surrounding area.
[0031] According to the above definition, the problem in the embodiments of the present application can be described as follows: given an arbitrary base station, the target is to develop a generative model Ψ that generates time series traffic data {xt}t=0:T under the condition of regional features {P, A} around the base station. It is not easy to solve the above problem. On the one hand, the generative model Ψ needs to explore the time characteristics of the mobile traffic sequence {xt}t=0:T, including periodic patterns and local fluctuations; on the other hand, the model also needs to capture the influence of regional factors {P, A} on mobile traffic. All these require the generative model to fully understand the complex dependence between the environment and network traffic.
[0032] As shown in Figure 1 , a structural schematic diagram of a Re-Diff framework provided by the embodiments of the present application is provided, which is composed of three main components: a time pattern extraction (TPE) unit, a regional environmental feature fusion (CFF) unit and an environment-aware scale learning (CSL) unit.
[0033] The time pattern extraction unit is used to capture the inherent time periodicity pattern in the traffic sequence, the regional environmental feature fusion unit extracts and fuses the regional features of the base station, especially the AOI and POI information. The environment-aware scale learning unit further combines these regional feature representations with the time features of the time pattern extraction unit to learn the relationship between the environment and the traffic sequence. Finally, the unconditional diffusion model generates the traffic sequence according to the regional environmental traffic sequence features learned by the environment-aware scale learning unit.
[0034] The geographic feature-oriented mobile traffic generation method provided by the embodiments of the present application will be described in detail below in combination with the drawings, specific embodiments and application scenarios.
[0035] As shown in Figure 2 , a geographic feature-oriented mobile traffic generation method provided by the embodiments of the present application, the method can include the following steps 201 and step 202: Step 201, obtaining an interest point vector and an interest area vector in the signal coverage range of a target base station, and extracting traffic sequence features from a randomly generated time domain traffic sequence, and extracting regional environmental features from the interest point vector and the interest area vector.
[0036] Among them, the interest point vector is used to represent the number of different types of interest points in the signal coverage range of the base station; the interest area vector is used to represent whether each interest area in the pre-divided multiple interest areas in the range of the base station exists; the traffic sequence feature is used to represent the traffic features of different frequency components.
[0037] Exemplarily, the time-domain traffic sequence is obtained based on real traffic data of the base station in the training process of the model, and is obtained based on randomly generated traffic data in the prediction process of the model. The target base station is any one of the following: a real base station at the target location, a virtual base station at the target location.
[0038] It can be understood that the mobile traffic generation method for geographical characteristics provided by the embodiments of the present application can simulate the local network traffic changes according to the local geographical environment characteristics in the area lacking of mobile historical traffic, and can effectively support a series of applications such as base station site selection, site planning, network hibernation of the network, and therefore the target base station can be a real base station or a virtual base station.
[0039] Specifically, the traffic sequence features include a plurality of frequency domain bases and a time domain residual sequence, and the step 201 of extracting the traffic sequence features from the randomly generated time-domain traffic sequence can further include the following steps 201a1 and 201a2: Step 201a1, converting the time-domain traffic sequence from time domain to frequency domain by using discrete Fourier transform to obtain a frequency-domain traffic sequence.
[0040] Step 201a2, dividing the frequency-domain traffic sequence into a plurality of frequency-domain bases and a time-domain residual sequence other than the plurality of frequency-domain bases.
[0041] Among them, the plurality of frequency-domain bases are used to represent the traffic characteristics of different frequency components, and one frequency-domain base represents the traffic characteristics of one frequency component.
[0042] Exemplarily, in order to capture the periodic characteristics of the time-domain traffic sequence, the time pattern extraction unit can convert the time-domain traffic sequence into the frequency domain and extract typical features. First, the original time-domain traffic sequence x with a length of L (i.e. the randomly generated time-domain traffic sequence) is converted into the frequency domain by using discrete Fourier transform (DFT) to obtain a frequency-domain traffic sequence X=F{x}.
[0043] Then, in order to find the frequency domain point set that contributes most to the periodicity, the time pattern extraction unit divides the sequence X into M frequency bases {Xm} and a time-domain residual sequence R. Let the index of the kth largest value in X be gk, where 0≤k<L, then Xm (containing N non-zero values) and R can be represented by the following formula one and formula two: (Formula one) (Formula two) Exemplarily, the above formula can be expressed as the mth N largest point to the [(m+1)N-1]th largest point in the frequency base Xm extraction X, and the other points are set to 0. Through the time mode extraction unit, the time domain sequence x can be converted and split into several time domain sequences xm representing different frequency domain bases, while the residual R is reserved to preserve the integrity of the sequence. Finally, the purpose of extracting different time period mode components and features is achieved.
[0044] Exemplarily, the regional environment feature fusion unit extracts and fuses two types of spatial information features, POI and AOI, by constructing a pre-trained traffic mean predictor. The main goal of the predictor is to model the relationship between regional attributes and traffic, that is, to learn an effective expression of traffic patterns from static spatial features.
[0045] Exemplarily, in terms of feature construction, first, for the location of each base station, the number of different types of POIs in the surrounding area (which can be within the base station signal coverage or within the preset range near the base station) is counted as a vector. Each dimension of the vector represents a POI type, such as catering, transportation facilities, commercial services, etc., and the corresponding numerical value represents the total number of interest points of that type. In this way, the functional density and use attributes of the area can be reflected.
[0046] In addition, to further enrich the spatial semantic information, AOI data, i.e., building use and land attribute information, is introduced. Similarly, an AOI vector is constructed for each base station, where each dimension represents the presence or absence of a certain type of AOI using one-hot encoding form of 0 or 1. This encoding method can effectively represent the functional division of the area, such as residential areas, commercial areas, industrial areas, etc.
[0047] Specifically, the step of extracting regional environment features from the POI vector and the AOI vector in step 201 can further include steps 201b1 and 201b2: Step 201b1, using a pre-trained traffic mean predictor to extract features from the POI vector and the AOI vector respectively, and fusing the extracted features based on an attention mechanism to obtain a traffic prediction mean.
[0048] Step 201b2, inputting the traffic prediction mean into a first multi-layer perception machine to obtain a corresponding traffic mean feature, and performing feature fusion on the traffic mean feature, the position encoding of the current time step, and the result feature obtained by predicting the POI vector based on a second multi-layer perception machine to obtain the regional environment feature.
[0049] Wherein, the position encoding is a sine position encoding generated based on the points on the time domain traffic sequence.
[0050] It can be understood that in time series data processing (such as traffic prediction, weather prediction), data is arranged in sequence according to time order (for example, hourly traffic, daily temperature). The "current time step" refers to the specific time that the model is processing.
[0051] Exemplarily, in order to make full use of these spatial structure information, a predictor is pre-trained to model the mapping relationship between POI and AOI features and traffic mean. In this way, the model can learn in advance which regional features will have a significant impact on traffic intensity. This not only provides more priori input for the subsequent main model, but also improves the model's ability to generalize in space.
[0052] It is worth mentioning that the predictor is not only used in the pre-training stage. During the training process of the main model, it continues to participate in loss optimization and is constantly updated with the training of the overall model, and its loss function is L1. This "training and adjustment" mechanism enables the predictor to always coordinate with the main model, thereby avoiding the problem of invalidating the early learning results in the subsequent stage.
[0053] As shown in Figure 1 In the main model, two independent multi-layer perception (MLP) modules are further designed by the embodiments of the application to process the predicted traffic mean and the original POI vector respectively. This separate processing method helps to capture different semantic information contained in the two types of input features. Next, the output results of the two MLPs are fused with the sine position encoding of the current time step to obtain the final regional environment embedding representation E (i.e. the above regional environment feature).
[0054] Exemplarily, through such a design, not only is the spatial structure information (POI, AOI) effectively integrated with the dynamic information (time), but also the learning efficiency and prediction accuracy of the model are improved through the pre-training mechanism. This is of great significance for modeling different regional traffic sequence scales, especially when facing data sparse or cold start regions, which can significantly enhance the adaptability and generalization ability of the model.
[0055] Step 202, input the regional environment feature as the key and value, and the traffic sequence feature as the query into the cross-attention module, calculate the corresponding to-be-processed sequence feature, and perform multiple denoising operations on the to-be-processed sequence feature to generate a mobile traffic sequence.
[0056] Exemplarily, after obtaining the traffic sequence feature extracted by the time pattern extraction unit and the regional environment feature learned by the regional environment feature fusion unit, the environmental perception scale learning unit can be used to explore the interaction relationship between the patterns of different frequency bases in the traffic sequence and the regional environment features in which they are located.
[0057] Exemplarily, consider the time series feature set output by the temporal pattern extraction unit, i.e. a series of flow feature representations containing different frequency components. It is desirable to further explore the connection between these different frequency components and the environmental embedding vectors. To this end, a cross-attention mechanism is introduced to establish a nonlinear, learnable association between the time-domain features and the spatial environment. In this mechanism, the regional environment embedding representation E is used as the key (Key) and value (Value), while the sequence xm output by the temporal pattern extraction unit is used as the query (Query). By taking the flow sequence features as the active party to "query" the environmental representation most relevant to it, the time and spatial dimensions can be more accurately aligned in the feature space. The output of this cross-attention module is represented as Tm, which reflects the matching relationship between the mth frequency component and the environmental features. This module can dynamically adjust the weight of each frequency component in the final representation, highlighting those components that are highly related to environmental factors and suppressing irrelevant or noise information.
[0058] Specifically, in the step 202, the step of calculating the corresponding to-be-processed feature sequence by inputting the regional environmental features as the key and the value, and the flow sequence features as the query into the cross-attention module, can further include the following steps 202a1 and 202a2: Step 202a1, cross-attention calculation is performed in the cross-attention module by taking the regional environmental features as the key and the value, and each frequency domain basis in the plurality of frequency domain bases as the query, to obtain a first feature corresponding to each frequency domain basis.
[0059] Step 202a2, after multiplying each frequency domain basis corresponding first feature with a corresponding parameter vector, a second feature corresponding to each frequency domain basis is obtained, and each frequency domain basis corresponding second feature is fused and added to the time-domain residual sequence to obtain the to-be-processed feature sequence.
[0060] Exemplarily, after completing the cross-attention operation, a set of potential features {Tm} representing the association between different frequency components and the environment, and a residual time-domain flow sequence R after removing the main frequency through the temporal pattern extraction module are obtained.
[0061] To integrate this information into the denoising network, the original time-domain flow sequence x can be subjected to multi-layer processing such as diffusion embedding and time embedding, and further fused with the time-domain residual sequence R and all cross-attention outputs Tm. In the fusion process, a learnable weight is assigned to the attention representation Tm of each frequency level, allowing the model to integrate multi-scale information in a nonlinear manner. This mechanism not only enhances the expression ability of the model, but also improves its adaptability to different flow patterns.
[0062] Exemplarily, the fusion result encodes the complex dependence relationship between multi-frequency, multi-scale and environmental variables while maintaining the original time-domain amplitude characteristics. This design enables the model to effectively model the way in which the flow is affected by regional attributes at different time periods, thereby improving the overall prediction accuracy and interpretability. The loss function of the final model is L = L1 + L2, where L2 is the MSE loss function of the denoising network, L1 + L2, where L2 is the MSE loss function of the denoising network, is a hyperparameter.
[0063] Specifically, the step 202 of performing multiple denoising operations on the to-be-processed sequence features to generate a mobile flow sequence can further include the following step 202b: Step 202b, performing multiple denoising operations on the to-be-processed sequence features by using the residual connection layer in the denoising network to obtain the mobile flow sequence.
[0064] Exemplarily, in order to more clearly illustrate the technical solutions in the embodiment, the training process of the mobile flow generation method facing geographical characteristics in the embodiment will be explained in detail as follows: 1.1, Collecting base station flow sequences and cleaning, deleting base stations with all 0 flow data values on a certain day, obtaining B base station length one week, granularity 1h flow sequence x(B, L), where L = 24 * 7 is the time length.
[0065] 1.2, Collecting POI distribution around the base station and AOI coverage area. Select all POIs within a 500-meter radius around each base station, and count by category to form a POI vector P(B, K), where K is the total number of POI categories. At the same time, the AOI one-hot encoding A(B, J) is directly determined according to the AOI area to which the base station belongs, where J is the total number of AOI categories.
[0066] 1.3, Setting hyperparameters and random seeds. The hyperparameters include: the number of components extracted in the frequency domain M, the coefficient of L1 in the loss function It is recommended to select M = 4 or 5, and λ = 0.1.
[0067] 2.1, Dividing the original data x, P, A into training set and test set, and normalizing according to the training set data.
[0068] 2.2.1, Input the training set x, P, A into the Re-Diff model as shown in Figure 1 The input of the TPE module is x(B, L) in the new dimension C after input convolution, and the output xm(B, L, C) is the extracted time-domain sequence, and R(B, L, C) is the time-domain residual sequence.
[0069] 2.2.2, The CFF module input is P(B, K) and A(B, J), both of which are outputted by the pre-trained traffic mean predictor to output the predicted mean μp(B, 1), which is then expanded to (B, L, 1) to obtain the predicted mean embedding of (B, L, C) through MLP1. And add the output d(B, L, C) with P(B, L, C) through MLP2 and time step embedding timeemb(B, L, C).
[0070] 2.2.3, The CSL module first uses the output d(B, L, C) of the CFF module as Key & Value, and uses the output xm(B, L, C) of the TPE module as Query, and does cross attention in the C dimension to get the output Tm(B, L, C), and then multiplies Tm(B, L, C) and trainable parameter vector (B, L, C) respectively and adds the time domain residual R(B, L, C) to become x'(B, L, C).
[0071] 2.3, The output x'(B, L, C) is connected through the residual connection layer in the diffusion model denoising network, and the generated traffic sequence is obtained step by step. The loss function value is obtained by using the generated traffic sequence, the real traffic sequence, the predicted traffic mean and the real traffic mean. Update the learnable parameters in the model by taking the gradient.
[0072] 3.1, Repeat 2.2-2.3 to get the trained model. Input the test set x, P, A to the trained model to test the effect.
[0073] For example, the overall effect and the comparison results of other algorithms are as shown in Figure 3 , wherein Time-GAN is a time series generation algorithm that can capture the time dependence in time data; 5GT-GAN is a model designed for real 5G traffic data generation; CSDI is only based on traffic data to generate traffic sequences; POI-Diff is a diffusion model that directly inputs the POI vector into the conditional information; KST-Diff is a knowledge-enhanced spatio-temporal data diffusion model using city knowledge graph. As shown in Figure 3 , the bold indicates the best performance, and the underlined indicates the second best performance. Each column Δ represents the percentage improvement of the Re-Diff model, and the last column is the average improvement.
[0074] The experimental results show that compared with the GAN-based method, the diffusion model provides better generation stability and accuracy, and demonstrates its potential in the mobile traffic generation task. In addition, as shown in Figure 4 , the explicit modeling of regional characteristics in this method significantly improves the model's transferability ( Figure 4The accuracy is increased by more than 14.57%. This encourages future research to explore more diverse regional environmental features (such as geographical attributes, population density) to further reveal the universal dependence between network usage and the environment, thereby improving model performance.
[0075] The method for generating mobile traffic based on geographical features provided in the embodiments of the present application first acquires a point-of-interest vector and an interest vector in the signal coverage range of a target base station, extracts traffic sequence features from a randomly generated time-domain traffic sequence, and extracts regional environmental features from the point-of-interest vector and the interest vector. Then, the regional environmental features are taken as keys and values, and the traffic sequence features are taken as a query to be input into a cross-attention module to calculate corresponding to-be-processed sequence features, and the to-be-processed sequence features are subjected to multiple denoising operations to generate a mobile traffic sequence. The point-of-interest vector is used to represent the number of different types of points of interest in the signal coverage range of the base station. The interest vector is used to represent whether each interest surface in a plurality of pre-divided interest surfaces exists in the range of the base station. The traffic sequence features are used to represent traffic features of different frequency components. In this way, by using a modeling method driven by geographical features and a conditional generation mechanism to adaptively capture the dependence between geographical environment and mobile traffic, the complex network traffic data can be modeled more accurately.
[0076] It should be noted that the method for generating mobile traffic based on geographical features provided in the embodiments of the present application can be executed by a device for generating mobile traffic based on geographical features, or a control module in the device for generating mobile traffic based on geographical features for executing the method for generating mobile traffic based on geographical features. In the embodiments of the present application, the device for generating mobile traffic based on geographical features is taken as an example to illustrate the device for generating mobile traffic based on geographical features provided in the embodiments of the present application.
[0077] It should be noted that the method for generating mobile traffic based on geographical features shown in each of the above method diagrams is illustratively described by taking one of the diagrams in the embodiments of the present application as an example. In specific implementation, the method for generating mobile traffic based on geographical features shown in each of the above method diagrams can also be implemented in combination with any other diagram that can be combined as illustrated in the above embodiments, which will not be described herein again.
[0078] The device for generating mobile traffic based on geographical features provided in the present application is described below. The device for generating mobile traffic based on geographical features described below can be correspondingly referred to the method for generating mobile traffic based on geographical features described above.
[0079] Figure 5A structure schematic diagram of a mobile traffic generation device for geographical characteristics provided by an embodiment of the present application is shown in Figure 5 Specifically, the structure schematic diagram comprises: The acquisition module 501 is configured to acquire a point of interest vector and an interest area vector in a signal coverage range of a target base station; the feature extraction module 502 is configured to extract a traffic sequence feature from a randomly generated time domain traffic sequence, and extract a regional environment feature from the point of interest vector and the interest area vector; and the traffic generation module 503 is configured to input the regional environment feature as a key and a value, and the traffic sequence feature as a query into a cross-attention module, calculate a corresponding to-be-processed sequence feature, and perform multiple denoising operations on the to-be-processed sequence feature to generate a mobile traffic sequence; wherein the point of interest vector is used to represent the number of different types of points of interest in the signal coverage range of the base station; the interest area vector is used to represent whether each interest area in a plurality of pre-divided interest areas in the range of the base station exists; and the traffic sequence feature is used to represent the traffic features of different frequency components.
[0080] Optionally, the traffic sequence feature comprises a plurality of frequency domain bases and a time domain residual sequence; the feature extraction module 502 is specifically configured to convert the time domain traffic sequence from a time domain to a frequency domain by using a discrete Fourier transform to obtain a frequency domain traffic sequence; and the feature extraction module 502 is specifically further configured to divide the frequency domain traffic sequence into a plurality of frequency domain bases and a time domain residual sequence other than the plurality of frequency domain bases; wherein the plurality of frequency domain bases are used to represent the traffic features of different frequency components, and one frequency domain base represents the traffic features of one frequency component.
[0081] Optionally, the feature extraction module 502 is specifically configured to perform feature extraction on the point of interest vector and the interest area vector respectively by using a pre-trained traffic mean prediction period, and fuse the extracted features based on an attention mechanism to obtain a traffic prediction mean; and the feature extraction module 502 is specifically further configured to input the traffic prediction mean into a first multi-layer perceptron to obtain a corresponding traffic mean feature, and perform feature fusion on the traffic mean feature, a position encoding of a current time step, and a result feature obtained by predicting the point of interest vector based on a second multi-layer perceptron to obtain the regional environment feature; wherein the position encoding is a sine position encoding generated based on a point on the time domain traffic sequence.
[0082] Optionally, the traffic generation module 503 is specific for performing cross-attention calculation on the regional environment feature as a key and a value and each of the plurality of frequency domain bases as a query in a cross-attention module to obtain a first feature corresponding to each of the plurality of frequency domain bases; and the traffic generation module 503 is specific for multiplying the first feature corresponding to each of the plurality of frequency domain bases with a corresponding parameter vector to obtain a second feature corresponding to each of the plurality of frequency domain bases, and adding the second feature corresponding to each of the plurality of frequency domain bases after fusion to the time domain residual sequence to obtain the to-be-processed feature sequence.
[0083] Optionally, the traffic generation module 503 is specific for performing multiple denoising operations on the to-be-processed sequence feature by using a residual connection layer in a denoising network to obtain the mobile traffic sequence.
[0084] The geographic feature-oriented mobile traffic generation apparatus provided in the application first acquires a point-of-interest vector and an interest area vector in a signal coverage range of a target base station, extracts a traffic sequence feature from a randomly generated time domain traffic sequence, and extracts a regional environment feature from the point-of-interest vector and the interest area vector; then inputs the regional environment feature as a key and a value and the traffic sequence feature as a query into a cross-attention module to calculate a corresponding to-be-processed sequence feature, and performs multiple denoising operations on the to-be-processed sequence feature to generate a mobile traffic sequence; wherein the point-of-interest vector is used to represent the number of different types of points of interest in the signal coverage range of the base station; the interest area vector is used to represent whether each of a plurality of pre-divided interest areas in the range of the base station exists; and the traffic sequence feature is used to represent traffic features of different frequency components. In this way, by using a geographic feature-driven modeling method and a conditional generation mechanism, the dependence between a geographic environment and mobile traffic can be adaptively captured to more accurately model complex network traffic data.
[0085] Figure 6 An example of an entity structure diagram of an electronic device is shown in FIG. 1. Figure 6As shown, the electronic device can include a processor 610, a communications interface 620, a memory 630, and a communications bus 640, wherein the processor 610, the communications interface 620, and the memory 630 complete mutual communication through the communications bus 640. The processor 610 can invoke a logical instruction in the memory 630 to execute a mobile traffic generation method oriented to geographical characteristics, which includes: first, obtaining a point-of-interest vector and an interest vector in a signal coverage range of a target base station, and extracting a traffic sequence feature from a randomly generated time-domain traffic sequence, and extracting a regional environment feature from the point-of-interest vector and the interest vector; then, inputting the regional environment feature as a key and a value and the traffic sequence feature as a query into a cross-attention module to calculate a corresponding to-be-processed sequence feature, and performing multiple denoising operations on the to-be-processed sequence feature to generate a mobile traffic sequence; wherein the point-of-interest vector is used to represent the number of different types of points of interest in the signal coverage range of the base station; the interest vector is used to represent whether each interest surface in a plurality of pre-divided interest surfaces in the range of the base station exists; and the traffic sequence feature is used to represent the traffic features of different frequency components. In this way, by using a geographical feature driven modeling method and a conditional generation mechanism to adaptively capture the dependency relationship between the geographical environment and the mobile traffic, the complex network traffic data can be more accurately modeled.
[0086] In addition, the logical instructions in the memory 630 described above can be implemented in the form of a software functional unit and sold or used as an independent product, and can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute all or part of the steps of the method described in various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.
[0087] In another aspect, the present application also provides a computer program product, which comprises a computer program stored on a computer readable storage medium, and the computer program comprises program instructions, when the program instructions are executed by a computer, the computer can execute the geographic feature-oriented mobile traffic generation method provided by the above method, which comprises the following steps: first, obtaining a point of interest vector and an interest surface vector in the signal coverage range of a target base station, extracting traffic sequence features from a randomly generated time domain traffic sequence, and extracting regional environment features from the point of interest vector and the interest surface vector; then, inputting the regional environment features as keys and values and the traffic sequence features as queries into a cross-attention module to calculate corresponding to-be-processed sequence features, and performing multiple denoising operations on the to-be-processed sequence features to generate a mobile traffic sequence; wherein the point of interest vector is used to represent the number of different types of points of interest in the signal coverage range of the base station; the interest surface vector is used to represent whether each interest surface in the pre-divided multiple interest surfaces in the range of the base station exists; and the traffic sequence features are used to represent traffic features of different frequency components. In this way, by using the geographic feature-driven modeling method and the conditional generation mechanism, the dependence between the geographic environment and the mobile traffic can be adaptively captured, and the complex network traffic data can be more accurately modeled.
[0088] In another aspect, the present application also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the geographic feature-oriented mobile traffic generation method provided by the above method, which comprises the following steps: first, obtaining a point of interest vector and an interest surface vector in the signal coverage range of a target base station, extracting traffic sequence features from a randomly generated time domain traffic sequence, and extracting regional environment features from the point of interest vector and the interest surface vector; then, inputting the regional environment features as keys and values and the traffic sequence features as queries into a cross-attention module to calculate corresponding to-be-processed sequence features, and performing multiple denoising operations on the to-be-processed sequence features to generate a mobile traffic sequence; wherein the point of interest vector is used to represent the number of different types of points of interest in the signal coverage range of the base station; the interest surface vector is used to represent whether each interest surface in the pre-divided multiple interest surfaces in the range of the base station exists; and the traffic sequence features are used to represent traffic features of different frequency components. In this way, by using the geographic feature-driven modeling method and the conditional generation mechanism, the dependence between the geographic environment and the mobile traffic can be adaptively captured, and the complex network traffic data can be more accurately modeled.
[0089] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purposes of the embodiments according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0090] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and necessary universal hardware platforms, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of software products, and the computer software products can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and include a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the methods described in each embodiment or some parts of the embodiments.
[0091] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for generating mobile traffic based on geographical characteristics, characterized in that: include: Obtaining point of interest vectors and surface of interest vectors within the signal coverage of the target base station, extracting traffic sequence features from the randomly generated time-domain traffic sequence, and extracting regional environmental features from the point of interest vectors and surface of interest vectors; Input the regional environmental features as keys and values and the traffic sequence features as queries into the cross attention module, calculate the corresponding sequence features to be processed, and perform multiple denoising operations on the sequence features to be processed to generate a mobile traffic sequence; Among them, the point of interest vector is used to characterize the number of different types of points of interest within the signal coverage range of the base station; the surface of interest vector is used to characterize whether each of the multiple pre-divided surfaces of interest exists within the range of the base station; the traffic sequence feature is used to characterize the traffic characteristics of different frequency components.
2. The method according to claim 1, characterized in that The target base station is any one of the following: a real base station at the target location, a virtual base station at the target location.
3. The method according to claim 1 or 2, characterized in that The traffic sequence characteristics include: multiple frequency domain bases and time domain residual sequences; The extracting of traffic sequence features from the randomly generated time-domain traffic sequence includes: The time domain traffic sequence is converted from the time domain to the frequency domain using discrete Fourier transform to obtain a frequency domain traffic sequence; Dividing the frequency domain traffic sequence into a plurality of frequency domain bases and a time domain residual sequence other than the plurality of frequency domain bases; The multiple frequency domain bases are used to characterize the flow characteristics of different frequency components, and one frequency domain base characterizes the flow characteristics of one frequency component.
4. The method according to claim 3, characterized in that The extracting of regional environmental features from the interest point vector and the interest surface vector includes: Using the pre-trained traffic mean prediction period, feature extraction is performed on the interest point vector and the interest surface vector respectively, and the extracted features are fused based on the attention mechanism to obtain the traffic prediction mean; Inputting the traffic prediction mean into a first multi-layer perceptron to obtain a corresponding traffic mean feature, and fusing the traffic mean feature, the position code of the current time step, and the result feature obtained by predicting the interest point vector based on the second multi-layer perceptron to obtain the regional environment feature; The position code is a sinusoidal position code generated based on points on the time domain traffic sequence.
5. The method according to claim 4, characterized in that The regional environmental features are used as keys and values, and the traffic sequence features are used as queries to input into the cross attention module, and the corresponding feature sequence to be processed is calculated, including: Using the regional environmental feature as a key and a value, respectively, and performing a cross-attention calculation with each of the multiple frequency domain bases as a query in a cross-attention module, to obtain a first feature corresponding to each frequency domain base; The first feature corresponding to each frequency domain basis is multiplied by the corresponding parameter vector to obtain the second feature corresponding to each frequency domain basis, and the second feature corresponding to each frequency domain basis is fused and added to the time domain residual sequence to obtain the feature sequence to be processed.
6. The method according to claim 5, characterized in that The performing multiple denoising operations on the sequence features to be processed to generate a mobile traffic sequence includes: The residual connection layer in the denoising network is used to perform multiple denoising operations on the sequence features to be processed to obtain the mobile traffic sequence.
7. A mobile traffic generation device oriented to geographical characteristics, characterized in that: The device comprises: An acquisition module, configured to acquire a point of interest vector and a plane of interest vector within the signal coverage range of a target base station; A feature extraction module is used to extract traffic sequence features from the randomly generated time-domain traffic sequence, and to extract regional environmental features from the interest point vector and interest surface vector; a traffic generation module, configured to input the regional environmental features as keys and values and the traffic sequence features as queries into a cross-attention module, calculate corresponding sequence features to be processed, and perform multiple denoising operations on the sequence features to be processed to generate a mobile traffic sequence; Among them, the point of interest vector is used to characterize the number of different types of points of interest within the signal coverage range of the base station; the surface of interest vector is used to characterize whether each of the multiple pre-divided surfaces of interest exists within the range of the base station; the traffic sequence feature is used to characterize the traffic characteristics of different frequency components.
8. The device according to claim 7, characterized in that The traffic sequence characteristics include: multiple frequency domain bases and time domain residual sequences; The feature extraction module is specifically used to convert the time domain traffic sequence from the time domain to the frequency domain using discrete Fourier transform to obtain a frequency domain traffic sequence; The feature extraction module is further configured to divide the frequency domain traffic sequence into a plurality of frequency domain bases and a time domain residual sequence other than the plurality of frequency domain bases; The multiple frequency domain bases are used to characterize the flow characteristics of different frequency components, and one frequency domain base characterizes the flow characteristics of one frequency component.
9. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps of the method for generating mobile traffic oriented to geographical characteristics as claimed in any one of claims 1 to 6 are implemented.
10. A computer-readable storage medium, characterized in that A computer program is stored thereon, and when the computer program is executed by a processor, the steps of the method for generating mobile traffic oriented to geographical characteristics as claimed in any one of claims 1 to 6 are implemented.