Multi-dimensional Traffic Data Synthesis Method and System Based on Large Language Model
Through the multi-dimensional traffic data synthesis method based on large language model, the problem that the prior art cannot synthesize heterogeneous data containing time series is solved, and high fidelity and high efficiency cellular traffic data synthesis is achieved.
Patent Information
- Application Number
- CN202410930622.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-11
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-07-11
AI Technical Summary
Existing cellular data synthesis methods cannot synthesize effective heterogeneous data containing time series under any conditions.
A multi-dimensional traffic data synthesis method based on a large language model is adopted to obtain historical table data for text feature encoding, an autoregressive large language model is constructed, and the model is fine-tuned using historical text encoding features to finally generate synthetic table data.
It realizes the flexible and effective synthesis of mobile cellular network data with high fidelity, high practicality and effective privacy protection under any conditions, and improves the efficiency of data synthesis.
Smart Images

Figure CN118886397B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of cellular data synthesis, and in particular, to a multi-dimensional traffic data synthesis method and system based on a large language model. Background Art
[0002] Cellular traffic data synthesis is an important research direction in the field of mobile communications. Existing research methods mainly include methods based on generative adversarial networks, methods based on diffusion models, and methods based on deep language models. The research based on generative adversarial networks and diffusion models lacks flexibility and cannot support conditional generation with specific combinations of features. In addition, data preprocessing operations (such as one-hot encoding of categorical data) are vulnerable to multi-dimensional data modeling, and they may also lose some important information due to data transformation (i.e., data scaling and smoothing). With the rapid development of large language models, thanks to the rich and efficient understanding and expression capabilities of large language models, it provides an opportunity for us to flexibly generate synthetic cellular traffic data with arbitrary feature conditions.
[0003] Cellular traffic data usually contains information in different dimensions, including user attributes and application (app) preferences, traffic sequences, etc. These data are collected from mobile phone users who use mobile applications on smartphones. Compared with existing methods based on GANs and diffusion models, large language models have two main advantages in traffic synthesis in the way of natural language processing. First, the tokenization process of large language models only processes text and does not require prior specification of the data type of features. The completely text-based tokenization method can alleviate the problem of lossy preprocessing and solve the dimensional explosion problem caused by one-hot encoding of high-dimensional data. Second, both feature names and feature values are included in the text encoding, which retains more context information than ordinary data transformation and is beneficial for learning models to model heterogeneous data with various features. However, existing table synthesis methods based on language models ignore the modeling of time series and cannot effectively generate cellular traffic sequences. Summary of the Invention
[0004] The present invention provides a multi-dimensional traffic data synthesis method and system based on a large language model, which is used to solve the technical problem that existing cellular data synthesis methods cannot synthesize effective heterogeneous data containing time series under arbitrary conditions.
[0005] To solve the above technical problem, the technical solution proposed by the present invention is as follows:
[0006] A multi-dimensional traffic data synthesis method based on a large language model, comprising:
[0007] Obtain historical tabular data, where the historical tabular data includes traffic-related data in multiple different dimensions;
[0008] Perform text feature encoding on the traffic-related data in the historical table data to obtain historical text encoding features;
[0009] Construct an autoregressive large language model and fine-tune the autoregressive large language model using the historical text encoding features;
[0010] Obtain target table data and sample the target table data using the autoregressive large language model, and generate synthetic table data according to the sampling results.
[0011] Preferably, the traffic-related data includes user attribute data, program usage time series data, and program usage traffic series data;
[0012] Performing text feature encoding on the traffic-related data in the historical table data includes:
[0013] Obtain the user attribute data in the historical table data, and perform text encoding on the user attributes to obtain user attribute features;
[0014] Obtain the program usage time series data in the historical table data, and perform text encoding on the time series data to obtain time series features;
[0015] Obtain the program usage traffic series data in the historical table data, and perform text encoding on the traffic series data to obtain traffic series features;
[0016] Concatenate the user attribute features, time series features, and traffic series features to obtain historical text encoding features.
[0017] Preferably,
[0018] The representation form of the user attribute feature encoding is:
[0019] where i represents the subscript of the user attribute, and j represents the subscript of the user attribute value; represents the feature name of the i-th user attribute, represents the feature value of the i-th user attribute; the superscript u indicates that both the feature name and the feature value come from the user attribute data;
[0020] and / or
[0021] The template for the time series feature encoding is: "Start tag - Time - End tag:
[0022] The start time series encoding is represented as The end time series encoding is represented as where \(i\) represents the subscript of the time feature, and \(j\) represents the subscript of the value of the time feature; \(TIME\_START\) and \( / TIME\_START\) are the prefix and suffix for marking the start time information of the sequence respectively, \(TIME\_END\) and \( / TIME\_END\) are the prefix and suffix for marking the end time information of the sequence respectively, \(f\) i d represents the feature name of the \(i\)-th time feature, represents the feature value of the \(i\)-th time feature, and the superscript \(d\) indicates that both the feature name and the feature value are from the time series data;
[0023] and / or
[0024] The representation form of the traffic sequence feature encoding is:
[0025] where \(i\) represents the subscript of the traffic sequence feature, and \(j\) represents the subscript of the traffic sequence feature value, where, is the feature name of the \(i\)-th traffic sequence feature, is the feature value of the \(i\)-th traffic sequence feature, and the superscript \(r\) indicates that both the feature name and the feature value are from the traffic sequence data;
[0026] and / or
[0027] The historical text encoding feature is \(t\) i,j =\(\{u\) i,j ,d i,j,start ,d i,j,end ,r i,j \}\), \(t\) i,j represents the \(j\)-th feature text encoding in the \(i\)-th consecutive time period.
[0028] Preferably, using the historical text encoding feature to fine-tune the autoregressive large language model includes:
[0029] Based on the historical table data, construct multiple historical text encoding features for consecutive time periods of the same user ID, and construct the time series text data \(t\) i =(t i,1 , ",", t i,2 , …, ",", t i,m ), where \(t\) i,1 is the first feature text encoding in the \(i\)-th consecutive time period; \(t\) i,2 is the second feature text encoding in the \(i\)-th consecutive time period; \(t\) i,m is the \(m\)-th feature text encoding in the \(i\)-th consecutive time period; \(m\) is the number of historical text encoding features;
[0030] Use the random permutation function for \(t\) iRearrange to obtain the sorted time series text data where [k1, k2, …, k m = P([1, 2, …, m]), where t i,1 is the first feature text encoding in the i-th consecutive time period; t i,2 is the second feature text encoding in the i-th consecutive time period; t i,m is the m-th feature text encoding in the i-th consecutive time period;
[0031] Decompose the sorted time series text data t′ i into single words, sub-words, and characters to obtain the token sequence [w1, w2, …, w j , w1 is the first token, w2 is the second token, w j is the j-th token;
[0032] Decompose and learn the probability of the token sequence through autoregression, and adjust the autoregressive large language model parameters;
[0033] Preferably, decompose and learn the probability of the token sequence through autoregression, and adjust the autoregressive large language model parameters, including:
[0034] Construct a loss function and an optimizer to train the autoregressive large language model;
[0035] The calculation method of the loss function is as follows:
[0036] where w k is the token that should be generated currently in the original token sequence, p(w k ) is the probability that the currently generated token is the token w k that should be generated currently in autoregressive training;
[0037] The process of the optimizer updating the weight parameters is as follows:
[0038] Calculate the first moment estimate value m k of the k-th gradient:
[0039] m k = β1m k-1 + (1 - β1)g k ; where k is the number of iterations, g k represents the parameter gradient of the k-th time, β1 is the first-order weight decay coefficient, and m k-1 is the first moment estimate value of the (k - 1)-th gradient;
[0040] Based on the first moment estimate value m k of the k-th gradient, calculate the second moment estimate value v of the gradientk :
[0041] β2 is the second - order weight decay coefficient, and v k-1 is the second - moment estimate of the (k - 1)-th gradient;
[0042] Correct the bias of the first - moment estimate:
[0043]
[0044] where is the corrected first - moment estimate;
[0045] Correct the bias of the second - moment estimate where is the corrected second - moment estimate;
[0046] Update the parameters using the following formula:
[0047]
[0048] where θ k represents the model parameters at the k - th time, θ k-1 represents the model parameters at the (k - 1)-th time, α is the learning rate, λ is the L2 regularization coefficient,
[0049] ∈ is a small constant added for numerical stability.
[0050] Preferably, decompose the probability of the learning token sequence by autoregressive means, which is achieved by the following formula:
[0051]
[0052] where t represents the time - series text sentence of the current training; p(t) represents the probability that the autoregressively generated text is exactly the original text; p(t) is the training objective, i.e., to maximize the probability.
[0053] Preferably, sample the target tabular data using the autoregressive large - language model, and generate synthetic tabular data according to the sampling results, including:
[0054] Pre - process the input prompt, and the prompt is decomposed into tokens [w1, w2, …, w k-1 , and the autoregressive large - language model predicts the next token w k according to the observed first k - 1 tokens until the end token is predicted;
[0055] Convert the token numbers predicted by the large language model into text, convert the generated text data into tabular data using predefined rules, and assign corresponding data types to the synthetic data according to the information of the feature columns in the original data;
[0056] Check the type and range of the generated data. If the generated feature data cannot be converted into the corresponding data type of the original data or the generated feature data exceeds the range defined by the original data, discard this part of the feature data and use the successfully synthesized data to reconstruct the feature key-value pair prompt for data regeneration operations until the types and ranges of all synthesized data are correct.
[0057] Preferably, the input prompt includes: a feature name prompt, and the input format of the feature name prompt is "[Feature]is", and the autoregressive large language model samples and generates samples from the joint distribution p(v1,…,v n ) where v i represents the random variable corresponding to feature i, where i ∈ {1,…,n}.
[0058] Preferably, the input prompt includes: a feature key-value prompt, and the input format of the feature key-value prompt is "[Feature i is[Values i ", where i ∈ {1,…,n}, and the autoregressive large language model samples and generates data from the distribution of the remaining features according to this prompt.
[0059] A computer system includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.
[0060] The present invention has the following beneficial effects:
[0061] 1. In the present invention, a text encoder is constructed to convert the original tabular data into text data according to the type of features. The text data describes the features and feature values of the original tabular data in the form of natural language. An autoregressive large language model is constructed, and the autoregressive large language model is fine-tuned using the tabular data after text encoding, enabling it to fully learn and understand the data distribution of each feature, the dependence relationship between different features, and the time dependence relationship between historical traffic data. A synthetic data sampler is constructed, and the fine-tuned large language model is used to sample in the original tabular data to generate text that is similar to the distribution of the original tabular data and has good data privacy according to different prompt inputs, and the generated text data is converted into synthetic tabular data. The present invention can flexibly and effectively synthesize mobile cellular network data with high fidelity, high practicality, and effective privacy protection under any conditions using a large language model, and adopts a regeneration strategy when synthesizing data, greatly improving the efficiency of data synthesis.
[0062] In addition to the purposes, features, and advantages described above, the present invention has other purposes, features, and advantages. The following will refer to the accompanying drawings to further elaborate on the present invention in detail. BRIEF DESCRIPTION OF THE DRAWINGS
[0063] The accompanying drawings that form a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0064] Figure 1 It is a framework diagram of a cellular traffic data synthesis model (LLMCell) in a preferred embodiment of the present invention;
[0065] Figure 2 It is a CDF comparison diagram of the EMD scores of the uplink traffic synthetic data in a preferred embodiment of the present invention;
[0066] Figure 3 It is a CDF comparison diagram of the EMD scores of the downlink traffic synthetic data in a preferred embodiment of the present invention;
[0067] Figure 4 It is a CDF comparison diagram of the JSD scores of the synthetic data of user application preference tags in a preferred embodiment of the present invention;
[0068] Figure 5 It is a MAE performance comparison diagram of using LSTM to predict traffic for the overall data in a preferred embodiment of the present invention;
[0069] Figure 6 It is a RMSE performance comparison diagram of using LSTM to predict traffic for the overall data in a preferred embodiment of the present invention;
[0070] Figure 7 This is the MAE performance comparison chart of using ConvLSTM to predict traffic for the overall data in the preferred embodiment of the present invention;
[0071] Figure 8 This is the RMSE performance comparison chart of using ConvLSTM to predict traffic for the overall data in the preferred embodiment of the present invention;
[0072] Figure 9 This is the MAE performance comparison chart of using LSTM to predict traffic for data with entertainment application preferences in the preferred embodiment of the present invention;
[0073] Figure 10 This is the MAE performance comparison chart of using LSTM to predict traffic for data with social media application preferences in the preferred embodiment of the present invention. Detailed implementation manners
[0074] The embodiments of the present invention will be described in detail below with reference to the accompanying drawings. However, the present invention can be implemented in many different ways defined and covered by the claims.
[0075] Embodiment 1:
[0076] A multi-dimensional traffic data synthesis method based on a large language model is disclosed in this embodiment, including the following steps:
[0077] Construct a text encoder, and convert the original tabular data into text data according to the type of features. The text data describes the features and feature values of the original tabular data in the form of natural language.
[0078] Construct an autoregressive large language model, and use the tabular data after text encoding to fine-tune the autoregressive large language model, so that it fully learns and understands the data distribution of each feature, the dependence relationship between different features, and the time dependence relationship between historical traffic data.
[0079] Construct a synthetic data sampler, use the fine-tuned large language model to sample in the original tabular data, generate text similar to the distribution of the original tabular data and with good data privacy according to different prompt inputs, and convert the generated text data into synthetic tabular data.
[0080] In addition, in this embodiment, a computer system is also disclosed, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the above method are implemented.
[0081] The present invention can flexibly and effectively synthesize mobile cellular network data with high fidelity, high practicality, and effective privacy protection under any conditions by using a large language model, and adopts a regeneration strategy when synthesizing data, greatly improving the efficiency of data synthesis.
[0082] Embodiment 2:
[0083] Embodiment 2 is a preferred embodiment of Embodiment 1. The difference between it and Embodiment 1 lies in expanding the specific steps of the flexible and effective cellular traffic data synthesis method based on a large language model:
[0084] In this embodiment, a flexible and effective cellular traffic data synthesis method based on a large language model is disclosed. Through the text encoding mode and the autoregressive fine-tuning mechanism of the large language model, the temporal dependence and the dynamic relationship between features of mobile cellular network data are effectively captured to improve the accuracy of cellular traffic data synthesis under any given conditions.
[0085] In this embodiment, taking the mobile cellular traffic data synthesis of a large operator as an example, this method is applied to the mobile cellular traffic data synthesis. It specifically includes the following steps:
[0086] Step 1: Preprocessing of mobile cellular traffic data
[0087] The dataset of this example contains the mobile cellular network traffic data records of a large operator. By cleaning and preprocessing the dataset, the mobile cellular traffic usage data of 9,900 users is obtained. The statistical time interval of the traffic data is in hours, and the traffic data of each user includes the cellular traffic data used within one month from December 1, 2020 to December 30, 2020. Each piece of user data includes: user ID, user age, user gender, the type of application preferred by the user, the time when the cellular traffic data is generated, and the sizes of the uplink traffic and downlink traffic generated by the user in the past hour. The present invention uses the entire dataset for the autoregressive training of the large language model, and divides the consecutive 24 pieces of cellular traffic data of each user into a sample. The entire training set contains a total of 306,900 samples.
[0088] Step 2: Construct a text encoder and convert the original tabular data into a text representation
[0089] Since the input form of the large language model is text, it is necessary to encode the tabular data into a text representation with semantic information. As Figure 1 shown in "TextualEncoder" in, the text encoder constructed by the present invention has a total of three modules: a user attribute label encoding module, a date and time encoding module, and a traffic sequence encoding module, which respectively encode different parts of the cellular traffic data.
[0090] The user attribute label encoding module encodes the user attribute information in each sample according to the formula For encoding. In one sample, although there are 24 pieces of data, they all belong to the same user. Therefore, each user attribute has only one value. For example, if the user age corresponding to a certain sample is 22, it will be encoded into the text "age is22,"
[0091] The date and time encoding module encodes the start sampling time of each sample, that is, the time of the first piece of data, according to the formula For encoding; encodes the end sampling time of each sample, that is, the time of the last piece of data, according to the formula For encoding. For example, if the start sampling time of a certain cellular traffic data sample is 0:00 on December 2, 2020, and the end sampling time is 23:00 on December 2, 2020, the start time will be encoded into the text "TIME_START date is20201202,houris0 / TIME_START", and the end time will be encoded into the text "TIME_END date is20201202,hour is23 / TIME_END".
[0092] The traffic sequence encoding module encodes the traffic sequence of each sample according to the formula For encoding. In one sample, there are a total of 24 pieces of data, and each piece of data has a traffic value. The traffic values of the 24 pieces of data form a traffic sequence with a length of 24. For example, if the uplink traffic values of 24 pieces of data in a certain sample are 1.52, 2.66, …, 55.219 respectively, the traffic sequence feature of this sample will finally be encoded into the text "Uplink traffic series is(1.52,2.66,…,55.219)".
[0093] After the user attributes, date and time, and traffic sequence are encoded respectively, they are combined into the final time series text data at intervals of commas
[0094] Step 3: Fine-tune the large language model based on the text-encoded tabular data
[0095] The model base used in this implementation example is GPT2 after knowledge distillation, that is, DistilGPT-2, which has certain semantic understanding capabilities and is suitable for fine-tuning with specific data to adapt to downstream tasks
[0096] After the data text encoding is completed, each sample will form a time series text data t i =(t i,1 ,".",t i,2 ,".",…,".",t i,m ), where t i,j= {u i,j , d i,j,start , d i,j,end , r i,j}. However, in text encoding transfer, the features of tabular data are arranged in a fixed order, and the coupling degree of the appearance order before and after the features is high, which will affect the sampling diversity when the fine-tuned model synthesizes data and the arbitrariness of the data synthesis conditions. To ensure the relative independence of the features, the present invention uses a random permutation function to re-arrange t i to obtain where After the text data is re-arranged, it provides benefits for synthesizing data under any conditions during subsequent sampling.
[0097] The present invention fine-tunes the large language model in an autoregressive manner. During the fine-tuning process, first, a specific tokenization scheme is used to preprocess the sorted text sentences t into individual words, subwords, and characters using a text corpus. Thus, the sentence t is decomposed into a series of token numbers [w1, w2, …, w j . Then, the language model decomposes and calculates the probability of the token number sequence, and adjusts the model parameters through the error between the predicted sequence and the true sequence. The estimated value of the entire sequence is obtained by predicting the conditional probability of the next token on the premise of the already calculated tokens, as shown in formula (1):
[0098]
[0099] Step Four: Generate tabular data based on the fine-tuned large language model
[0100] After the fine-tuning step, a fine-tuned large language model is obtained. This model has learned the distribution of the same features of the original tabular data at different times and the correlation relationship between different features, and can sample from the original data to generate synthetic data similar to the distribution of the original data.
[0101] During the data generation process, for each sample to be generated, the large language model is given text prompts for the start time and end time of the sample, such as "TIME_S RT date is 20201202, hour is 0 / TIME_START, TIME_END date is 20201202, hour is 23 / TIME_END". The large language model converts the text prompt into tokens and predicts the probability of the next token w k-1 based on the observed tokens [w1, w2, …, w k , predicts the remaining feature information, and finally synthesizes a sample data.
[0102] After obtaining the sample data, corresponding features and feature values are extracted from the text data using specific regular expressions and converted into the same type as the original tabular data. Due to the existence of random noise in the data generation process, the generated data may be missing or exceed the range defined by the original data. Therefore, after data conversion, abnormal data is discarded, and the correctly generated synthetic data is used to reconstruct the feature key-value pair prompt. The large model uses the regenerative strategy to generate the missing data, rather than all the data, to improve the data synthesis efficiency. After all feature values are correctly synthesized, the data is saved as a CSV file to obtain the final tabular data.
[0103] Step Five: Model Performance Evaluation
[0104] To verify the effectiveness of the model of the present invention, fidelity, usability, and privacy evaluation experiments are conducted using synthetic data. The data synthesis method is shown in Steps 1 - 4.
[0105] In this step, the data synthesis model LLMCell in this embodiment is compared with the following several state-of-the-art data synthesis methods, including STAN, DoppelGANger, TabDDPM, and GReaT. Among them, STAN is a network traffic data synthesis tool using an autoregressive neural model. DoppelGANger is a GAN-based multi-dimensional time series generation model. TabDDPM is a model that uses a denoising diffusion probability model to generate synthetic tabular data sets. GReaT is a method that uses an autoregressive generative large model to sample synthetic tabular data.
[0106] Step 5.1 Fidelity Evaluation Experiment
[0107] In the data synthesis task, the synthetic data should be similar to the distribution of the original data. The degree of consistency between the synthetic data and the original data distribution is called fidelity. In this embodiment, the fidelity of the synthetic uplink and downlink traffic sequences is evaluated using Earth Mover Distance (EMD), and the classification features such as user application preference labels are evaluated using Jensen-Shannon divergence (JSD).
[0108] EMD is used to calculate the similarity of continuous type data in the real data and the synthetic data. Let P and Q represent two probability distributions of a given set X respectively. Then the calculation method of EMD is shown in formula (2), where d(x, y) represents the distance between elements x and y in the two probability distributions, and γ(x, y) represents the mass size transferred from each element x in P to each element y in Q under the premise of using the optimal transportation plan γ. The smaller the EMD value, the higher the similarity of the continuous type data in the real data and the synthetic data, and the higher the data fidelity.
[0109]
[0110] JSD is used to calculate the similarity of categorical data in real data and synthetic data. Let P and Q represent two probability distributions of a given set X respectively. Then the calculation method of JSD is shown in formula (3), where P(x) and Q(x) are the probabilities of x in distributions P and Q respectively. The smaller the JSD value, the higher the similarity of categorical data in real data and synthetic data, and the higher the data fidelity.
[0111]
[0112] The present invention uses baseline data synthesis methods such as STAN, DoppelGANger, TabDDPM, and GReaT to conduct a comparative evaluation of the fidelity with the LLMCell model. Among them, the synthetic data of 9,900 users is divided into 44 groups, with each group containing 225 users. The average EMD values of the uplink and downlink traffic of each group of data for different synthesis methods are calculated respectively, and the average JSD value of the user application preferences is calculated. The results are respectively normalized to the range of [0.1, 0.9], and the CDF graph is plotted. The CDF graph of the uplink traffic EMD score is as Figure 2 shown, the CDF graph of the downlink traffic EMD score is as Figure 3 shown, and the CDF graph of the user application preference JSD score is as Figure 4 shown. The following conclusions can be drawn: 1) The LLMCell model proposed by the present invention is significantly superior to other methods in terms of the fidelity of uplink and downlink traffic. Considering the 90th percentile, for the uplink traffic, the EMD scores of GreaT, STAN, DGAN, and TabDDPM are approximately equal to 0.15, 0.13, 0.62, and 0.63 respectively, while the EMD score of LLMCell is 0.10, which is improved by 30.12%, 20.75%, 83.54%, and 83.85% respectively; similarly, for the downlink traffic, the EMD scores of GreaT, STAN, DGAN, and TabDDPM are 0.18, 0.13, 0.55, and 0.58 respectively. The score of the LLMCell scheme is only 0.10. This proves the customized design of LLMCell, especially the text encoding of the traffic sequence, which fully considers the time series pattern while other baselines do not consider this factor carefully. 2) LLMCell is superior to STAN, DGAN, and TabDDPM methods in terms of the fidelity of user application preferences. In particular, at the 80th percentile, the JSD values of STAN, DGAN, and TabDDMP reach about 0.33, 0.77, and 0.77 respectively, while the JSD values of GreaT and the LLMCell model proposed by the present invention are only 0.03 and 0.07. This proves that the LLM-based method has a powerful inference ability based on text representation and can effectively model categorical data.
[0113] Step 5.2 Practicality Evaluation Experiment
[0114] The present invention evaluates the practicality of the proposed LLMCell model through two common downstream applications, namely uplink traffic prediction and downlink traffic prediction. The present invention uses the mean absolute error (MAE) and the root mean square error (RMSE) as evaluation metrics for uplink traffic and downlink traffic prediction. These two metrics are widely used in the evaluation of regression prediction tasks, and the smaller the metric value, the better. Let y i and represent the true value and the predicted value of the traffic at the i-th time step respectively, and n is the total number of predictions. The calculation methods are as shown in Formulas (4) and (5):
[0115]
[0116] The present invention adopts two common time series prediction models, including the long short-term memory network and the convolutional long short-term memory network.
[0117] Long Short-Term Memory Network (LSTM): A variant of the recurrent neural network (RNN), which is widely used in deep learning models for time series prediction.
[0118] Convolutional Long Short-Term Memory Network (ConvLSTM): A time series prediction model that combines the convolutional neural network (CNN) and the long short-term memory network (LSTM).
[0119] To test the generalization of the model trained on synthetic data, the present invention sorts the real data and the synthetic data based on the timestamp respectively, and then divides them into a 70% training set and a 30% test set respectively, and compares the metrics in two cases: 1) Train on real data and test on real data; 2) Train on synthetic data and test on real data.
[0120] (1) Overall Performance Comparison
[0121] To evaluate the overall performance of the LLMCell model, this embodiment compares the downstream applications for all synthetic data. Figure 5 and Figure 6 represent the average performance of the MAE metric with error bars and the LSTM metric evaluated using the LSTM predictor respectively, Figure 7 and Figure 8Respectively represent the average performance of the MAE metric and the LSTM metric with error bars evaluated using the ConvLSTM predictor. It can be observed that among the LSTM and ConvLSTM predictors, compared with other baselines, the LLMCell model has the best evaluation result performance and is closest to the real data. Specifically, when using the LSTM model for uplink traffic prediction, the MAE values of the synthetic data of GreaT, TabDDPM, DGAN, and STAN are 68.25, 71.91, 671.63, and 4261.39 respectively, and the corresponding RMSE values are 135.76, 150.81, 683.49, and 4264.13 respectively. In contrast, the MAE and RMSE values of LLMCell are only 59.44 and 122.78 respectively. In addition, when observing the error bars, LLMCell shows a lower deviation among all baselines, indicating the robustness of the synthetic data, further demonstrating that LLMCell can effectively capture the temporal patterns of cellular traffic data.
[0122] (2) Comparison of the performance of users with different application preferences
[0123] To evaluate the robustness of the data utility, the present invention further examines the downstream task performance of user data with different application preference labels. Figure 9 and 10 Represent the performance of the MAE metric evaluated using the LSTM predictor for two representative applications (entertainment and social media) with application preferences. It can be observed that LLMCell obtains the MAE score closest to the real data and is superior to other baselines, indicating that LLMCell still has the best performance under different application preference labels, further demonstrating the robustness of the synthetic data of the LLMCell model.
[0124] Step 5.3 Privacy evaluation experiment
[0125] To ensure the synthesis of high-quality cellular traffic data while preventing the leakage of user privacy, the present invention uses the closest distance record (DCR) metric to evaluate the distance-based privacy of each model. Let x and y represent the real data and the synthetic data respectively, and the calculation method of DCR is shown in formula (6). A larger DCR usually indicates lower data quality, while a smaller DCR may imply the potential risk of leaking sensitive information. Therefore, the privacy evaluation is relative, and the synthetic data model should maintain a relatively balanced position in the DCR.
[0126]
[0127] Table 1 shows the DCR metrics calculated from the uplink and downlink traffic data synthesized by different models and the real data. The results show that the synthetic data generated by TabDDPM has the shortest distance from the real data, which means that its synthetic data essentially mimics some real data points and has the highest possibility of privacy infringement; while the synthetic data generated by GreaT shows the largest distance, indicating that it generates new data rather than a copy close to the real data; by considering the comprehensive performance, the DCR achieved by LLMCell can balance data fidelity, utility, and privacy compared to the baseline.
[0128] Table 1 Results of Privacy Evaluation Experiment
[0129]
[0130]
[0131] In summary, the flexible and effective cellular traffic data synthesis method based on large language models disclosed in the present invention effectively captures the temporal dependence and dynamic relationship between features of mobile cellular network data through the text encoding mode and the autoregressive fine-tuning mechanism of the large language model, so as to improve the accuracy of cellular traffic data synthesis under any given conditions.
[0132] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A multi-dimensional traffic data synthesis method based on a large language model, characterized in that: include: Acquire historical table data, wherein the historical table data includes flow-related data of multiple different dimensions; Performing text feature encoding on the flow-related data in the historical table data to obtain historical text encoding features; A large autoregressive language model is constructed, and the large autoregressive language model is fine-tuned using the historical text encoding features; Acquire target table data, and sample the target table data using the autoregressive large language model, and generate synthetic table data according to the sampling results; Wherein, fine-tuning the autoregressive large language model using the historical text encoding features comprises: Based on the historical table data, multiple historical text encoding features of continuous time periods of the same user ID are constructed to construct time series text data t i =(t i,1 ,",",t i,2 ,…,",”,t i,m ), where t i,1 Encode the features of the first historical text in the i-th continuous time period; t i,2 Encode features for the second historical text in the i-th continuous time period; t i,m is the mth historical text encoding feature in the i-th continuous time period; m is the number of historical text encoding features; Use the random permutation function t i Rearrange to get sorted time series text data in in, Encode features for the first historical text in the i-th continuous time period after sorting; Encode features for the second historical text in the i-th continuous time period after sorting; Encode features for the mth historical text in the i-th continuous time period after sorting; For the sorted time series text data t′ i Decompose into single words, subwords and characters, and get the token sequence [w1,w2,…,w j ], w1 is the first mark number, w2 is the second mark number, w j is the jth marker number; The probability of learning the sequence of tokens is decomposed through autoregression, and the parameters of the autoregressive large language model are adjusted.
2. The multi-dimensional traffic data synthesis method based on a large language model according to claim 1 is characterized in that: The traffic-related data includes user attribute data, program usage time series data, and program usage traffic series data; Performing text feature encoding on the flow-related data in the historical table data includes: Acquire user attribute data in the historical table data, and perform text encoding on the user attributes to obtain user attribute features; Obtaining program usage time series data in the historical table data, and performing text encoding on the time series data to obtain time series features; Obtaining program usage flow sequence data in the historical table data, and performing text encoding on the flow sequence data to obtain flow sequence features; The user attribute features, time series features and flow series features are concatenated to obtain historical text coding features.
3. The multi-dimensional traffic data synthesis method based on a large language model according to claim 2 is characterized in that: The user attribute feature encoding is expressed in the form of: Where i represents the subscript of the user attribute, j represents the subscript of the user attribute value; f i u Represents the characteristic name of the i-th user attribute, represents the feature value of the i-th user attribute; the superscript u indicates that both the feature name and the feature value are derived from the user attribute data; and / or The template of the time series feature coding includes a start tag, a time and an end tag, which is divided into a start time series coding and an end time series coding; The start time sequence encoding is expressed as The end time series encoding is expressed as Where i represents the time feature subscript, j represents the subscript of the time feature value; TIME_START and / TIME_START are respectively used to mark the prefix and suffix of the sequence start time information, TIME_END and / TIME_END are respectively used to mark the prefix and suffix of the sequence end time information, and f i d Represents the feature name of the i-th time feature, It represents the characteristic value of the i-th time feature. The superscript d indicates that both the characteristic name and the characteristic value are from time series data. and / or The expression form of the traffic sequence feature coding is: Where i represents the subscript of the traffic sequence feature, j represents the subscript of the traffic sequence feature value, and f i r is the feature name of the i-th traffic sequence feature, is the eigenvalue of the ith traffic sequence feature, and the superscript r indicates that both the feature name and the eigenvalue are derived from the traffic sequence data; and / or The historical text encoding feature is t i,j = {u i,j ,d i,j,start ,d i,j,end ,r i,j }, t i,j Represents the jth feature text encoding in the i-th continuous time period.
4. The multi-dimensional traffic data synthesis method based on a large language model according to claim 3 is characterized in that: Decompose the probability of learning the sequence of tokens through autoregression and adjust the parameters of the autoregressive large language model, including: Constructing a loss function and an optimizer to train the autoregressive large language model; The loss function is calculated as follows: Among them, w k The current tag number to be generated in the original tag number sequence, p(w k ) is the autoregressive training, the current generated tag number is the current tag number w that should be generated k probability; The process of the optimizer updating weight parameters is as follows: Calculate the first-order moment estimate m of the k-th gradient k : m k =β1m k-1 +(1-β1)g k ; where k is the number of iterations, g k represents the kth parameter gradient, β1 is the first-order weight attenuation coefficient, m k-1 is the first-order moment estimate of the k-1th gradient; Calculate the second-order moment estimate v of the k-th gradient k : β2 is the second-order weight attenuation coefficient, v k-1 is the second-order moment estimate of the k-1th gradient; Correct the bias in the first moment estimate: in, is the corrected first-order moment estimate, is the k-th first-order weight attenuation coefficient; Correcting bias in second moment estimates in, is the corrected second-order moment estimate, is the k-th second-order weight attenuation coefficient; The parameters are updated using the following formula: Among them, θ k represents the k-th model parameter, θ k-1 represents the k-1th model parameter, α is the learning rate, λ is the L2 regularization coefficient, and ∈ is a small constant added for numerical stability.
5. The multi-dimensional traffic data synthesis method based on a large language model according to claim 4 is characterized in that: The probability of learning the sequence of labeled numbers is decomposed by autoregression, which is achieved by the following formula: Among them, t represents the time series text sentence of this training; p(t) represents the probability that the text generated by autoregression is exactly the original text.
6. The multi-dimensional traffic data synthesis method based on a large language model according to claim 5 is characterized in that: The target table data is sampled using the autoregressive large language model, and synthetic table data is generated according to the sampling result, including: The input prompt is preprocessed and decomposed into token numbers [w1,w2,…,w k-1 ], the autoregressive large language model predicts the next token number w based on the observed previous k-1 token numbers k , until the end tag number is predicted; Convert the token number predicted by the large language model into text, and convert the generated text data into feature names and feature values of synthetic table data using predefined rules, and then assign corresponding data types to the synthetic table data according to the feature name information in the historical table data; A type and range check is performed on the feature value of the feature name of the synthetic table data. If the feature value of the feature name of the synthetic table data cannot be converted into the data type corresponding to the historical table data or the feature value of the feature name of the synthetic table data exceeds the range specified by the historical table data, this part of the feature data is discarded, and other successfully synthesized feature values in the synthetic table data and their corresponding feature names are used to reconstruct feature key-value pair prompts, and data regeneration operations are performed until the feature value types and ranges of all feature names of the synthetic table data are correct.
7. The multi-dimensional traffic data synthesis method based on a large language model according to claim 6 is characterized in that: The input prompt includes: a feature name prompt, the input format of the feature name prompt is "[Feature]is", and the autoregressive large language model is based on the feature name prompt from the joint distribution p(v1,…,v n ) and generate samples, where v i represents the joint distribution p(v1,…,v n ), where i∈{1,…,n}, and n is the total number of random variables in the joint distribution.
8. The multi-dimensional traffic data synthesis method based on a large language model according to claim 6 is characterized in that: The input prompt includes: a feature key-value pair prompt, and the input format of the feature key-value pair prompt is "[Feature i ]is[Values i ]", the Feature i is the feature name of the i-th dimension in the feature key-value pair, Values i is the value of the feature name of the i-th dimension in the feature key-value pair, where i∈{1,…,n}. The autoregressive large language model generates data by sampling from the distribution of other feature key-value pairs based on this prompt.
9. A computer system comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.
Citation Information
Patent Citations
Audit report generation method and device, computer equipment and storage medium
CN117993369A
Cellular user App use data synthesis method based on large language model
CN118890612A