Method and apparatus for predicting network traffic, electronic device, and storage medium

CN120018194BActive Publication Date: 2026-09-18INSPUR TIANYUAN COMM INFORMATION SYST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510050718.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-13
Publication Date
2026-09-18
Estimated Expiration
2045-01-13

AI Technical Summary

Technical Problem

[0004]本发明提供一种网络业务量的预测方法、装置、电子设备及存储介质,用以解决现有技术中网络业务量的预测并不准确的缺陷,实现提高网络业务量预测的准确性

Benefits of technology

[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the network traffic prediction methods described above.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120018194B_ABST
    Figure CN120018194B_ABST
Patent Text Reader

Abstract

The application provides a network traffic prediction method, device, electronic equipment and storage medium, relates to the technical field of wireless networks, obtains multi-modal fusion features affecting network traffic according to multi-modal data, comprehensively collects features affecting network traffic of a region to be predicted, and is favorable for improving the accuracy of subsequent acquisition of traffic density maps and prediction of traffic data. Kernel density estimation is performed on the multi-modal fusion features according to a kernel function and a bandwidth parameter, traffic density maps are acquired, and accurate calculation of traffic density of the region to be predicted is realized. Through a traffic prediction model, feature extraction and feature analysis are automatically performed on the traffic density maps and the multi-modal fusion features, prediction traffic data is acquired, and the efficiency of acquisition of prediction traffic data is improved. The application improves the accuracy of acquisition of prediction traffic and improves the efficiency of acquisition of prediction traffic according to the traffic prediction model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wireless network technology, and in particular to a method, apparatus, electronic device, and storage medium for predicting network traffic. Background Technology

[0002] With the rapid development of wireless communication technology and the continuous increase in the number of mobile Internet users, the wireless network environment is becoming increasingly complex, and the volatility and unpredictability of traffic volume are also increasing. In this environment, traditional wireless network traffic prediction technologies often fail to meet the requirements of real-time performance and accuracy. These traditional technologies typically rely on a single data source (such as historical traffic data) or simple statistical models, making it difficult to accurately capture and analyze the complex dynamics of traffic volume changes.

[0003] In conclusion, the current predictions of network traffic volume are inaccurate. Summary of the Invention

[0004] This invention provides a method, apparatus, electronic device, and storage medium for predicting network traffic volume, in order to address the shortcomings of inaccurate network traffic volume prediction in the prior art and improve the accuracy of network traffic volume prediction.

[0005] This invention provides a method for predicting network traffic volume, comprising: extracting and fusing features from multimodal data affecting network traffic volume in the region to be predicted to obtain multimodal fused features; determining kernel function and bandwidth parameters based on the distribution pattern of the multimodal fused features; estimating kernel density of the multimodal fused features based on kernel function and bandwidth parameters to obtain a traffic volume density map; and inputting the traffic volume density map and multimodal fused features into a traffic volume prediction model to obtain predicted traffic volume data output by the traffic volume prediction model. The traffic volume prediction model is trained based on the labels of sample traffic volume density maps, sample multimodal fused features, and sample predicted traffic volume data, using a pre-defined deep learning model.

[0006] According to the network traffic prediction method provided by the present invention, the multimodal data is obtained based on the following steps: obtaining the geographical layout of the area to be predicted, the building density of the area to be predicted, and the traffic flow of the area to be predicted based on satellite imagery and urban monitoring systems to obtain video data of the area to be predicted; obtaining the environmental noise level of the area to be predicted at different times and locations based on urban noise monitoring systems to obtain audio data of the area to be predicted; obtaining text data affecting the communication demand of the area to be predicted based on publicly available data from public platforms; and obtaining multimodal data based on video data, audio data, and text data.

[0007] According to the network traffic prediction method provided by the present invention, feature extraction and feature fusion are performed on multimodal data affecting network traffic in the region to be predicted to obtain multimodal fused features. The method includes: preprocessing the multimodal data to obtain preprocessed video data, preprocessed audio data, and preprocessed text data; the preprocessing includes uniform formatting, data cleaning, and standardization; an image feature processing module based on a contrastive language-image pre-trained CLIP model extracts features from the preprocessed video data to obtain image spatial features; a text feature processing module based on the CLIP model extracts semantic and contextual features from the preprocessed text data to obtain event semantic features; a multilayer perceptron model extracts noise intensity features from the preprocessed audio data to obtain noise features; and the image spatial features, event semantic features, and noise features are fused to obtain the multimodal fused features.

[0008] According to the network traffic prediction method provided by the present invention, the kernel function and bandwidth parameters are determined based on the distribution pattern of multimodal fusion characteristics, including: when the distribution pattern is a normal distribution, the kernel function is determined to be a Gaussian kernel function, and the bandwidth parameters of the Gaussian kernel function are determined based on grid search and cross-validation; when the distribution pattern is a concentrated distribution, the kernel function is determined to be an Epanechnikov kernel function, and the bandwidth parameters of the Epanechnikov kernel function are determined based on grid search and cross-validation; when the distribution pattern is a long-tailed distribution, the kernel function is determined to be a hyperbolic tangent kernel function, and the bandwidth parameters of the hyperbolic tangent kernel function are determined based on grid search and cross-validation.

[0009] According to the network traffic prediction method provided by the present invention, the multimodal fusion feature includes multiple feature vectors. The kernel density of the multimodal fusion feature is estimated based on the kernel function and the bandwidth parameter to obtain a traffic density map. The method includes: obtaining the probability density estimate of each feature vector with the kernel function as the center and the bandwidth parameter as the radius; and obtaining the traffic density map based on the probability density estimates of all feature vectors.

[0010] According to the network traffic prediction method provided by the present invention, the traffic prediction model is determined based on the following steps: obtaining labels for the sample predicted traffic data based on the sample traffic heatmap, the sample predicted traffic value, and the sample traffic peak time period; labeling the sample traffic density map and the sample multimodal fusion features based on the labels of the sample predicted traffic data to obtain labeled training samples; training, testing, and validating a preset deep learning model based on the training samples until the error output by the preset deep learning model is less than a set error, thereby obtaining the traffic prediction model.

[0011] According to the network traffic prediction method provided by the present invention, the bandwidth parameter of the Gaussian kernel function is determined based on grid search and cross-validation, including: obtaining multiple initial bandwidth parameters based on the numerical range of multimodal fusion features; taking each initial bandwidth parameter as a grid, performing grid search and cross-validation on each grid according to the multimodal fusion features and the Gaussian kernel function to obtain an evaluation value for each initial bandwidth parameter; and taking the initial bandwidth parameter with the highest evaluation value as the bandwidth parameter.

[0012] This invention also provides a network traffic prediction device, comprising: a feature extraction module for extracting and fusing features from multimodal data affecting network traffic in the region to be predicted, to obtain multimodal fused features; a determination module for determining a kernel function and bandwidth parameters based on the distribution pattern of the multimodal fused features; a traffic density map acquisition module for estimating the kernel density of the multimodal fused features based on the kernel function and bandwidth parameters, to obtain a traffic density map; and a prediction module for inputting the traffic density map and the multimodal fused features into a traffic prediction model, to obtain predicted traffic data output by the traffic prediction model, wherein the traffic prediction model is trained based on the labels of sample traffic density maps, sample multimodal fused features, and sample predicted traffic data, on the basis of a preset deep learning model.

[0013] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement any of the network traffic prediction methods described above.

[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements any of the network traffic prediction methods described above.

[0015] The network traffic prediction method, apparatus, electronic device, and storage medium provided by this invention acquire multimodal fusion features affecting network traffic based on multimodal data, achieving comprehensive collection of features influencing network traffic in the area to be predicted, which is beneficial to improving the accuracy of subsequent acquisition of traffic density maps and predicted traffic data. Kernel density estimation is performed on the multimodal fusion features based on kernel functions and bandwidth parameters to obtain a traffic density map, achieving accurate calculation of traffic density in the area to be predicted. Through a traffic prediction model, automatic feature extraction and feature analysis are performed on the traffic density map and multimodal fusion features to obtain predicted traffic data, improving the efficiency of obtaining predicted traffic data. This invention improves the accuracy of obtaining predicted traffic based on multimodal fusion features and kernel density estimation, and improves the efficiency of obtaining predicted traffic based on a traffic prediction model. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0017] Figure 1 This is one of the flowcharts illustrating the network traffic prediction method provided by the present invention.

[0018] Figure 2 This is the second flowchart of the network traffic prediction method provided by the present invention.

[0019] Figure 3 This is a schematic diagram of the network traffic prediction device provided by the present invention.

[0020] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation

[0021] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0022] The following is combined with Figures 1-4 The present invention describes a method, apparatus, and electronic device for predicting network traffic.

[0023] Figure 1 This is one of the flowcharts illustrating the network traffic prediction method provided by the present invention, such as... Figure 1 As shown, the network traffic prediction method includes steps S100 to S400, and the specific steps are as follows.

[0024] S100: Perform feature extraction and feature fusion on the multimodal data affecting network traffic in the region to be predicted to obtain multimodal fused features.

[0025] Multimodal data refers to data with multiple modalities, such as video data, audio data, and text data. This multimodal data can capture different types of factors that affect the network traffic volume of a region, such as video data (capturing factors such as urban layout and traffic flow), audio data (capturing factors such as urban noise levels), and text data (capturing factors such as event and activity information).

[0026] Each modality's data source reflects factors directly related to network traffic volume. For example, video data reflects the density and activity in the area; audio data reflects the intensity of urban activity through noise levels; and text data provides information about events that may affect communication demand. These factors all influence the wireless network traffic volume in the area. Therefore, by extracting and fusing features from multimodal data, multimodal fusion features can be obtained, which can then more comprehensively predict traffic volume.

[0027] S200: Determine the kernel function and bandwidth parameters based on the distribution pattern of multimodal fusion features.

[0028] Multimodal fusion features consist of multiple feature vectors, each constituting a feature point. In kernel density estimation, the kernel function is used to calculate the density near the feature point. The bandwidth parameter defines the smoothness of the kernel density estimation, i.e., its sensitivity to local regions. A larger bandwidth parameter results in a smoother estimation, suitable for analyzing macroscopic trends; a smaller bandwidth parameter results in a more refined estimation, better suited for capturing microscopic details.

[0029] Based on the distribution pattern and numerical range of the multimodal fusion features in the region to be predicted, the kernel function and bandwidth parameters are determined to ensure the accuracy of kernel density estimation of the multimodal fusion features based on the kernel function and bandwidth parameters.

[0030] S300: Based on kernel function and bandwidth parameters, perform kernel density estimation on multimodal fusion features to obtain a traffic density map.

[0031] The multimodal fusion feature includes multiple feature vectors. Based on the kernel function and bandwidth parameter, the kernel density of the multimodal fusion feature is estimated to obtain the traffic density map. Specifically, with the kernel function as the center and the bandwidth parameter as the radius, the probability density estimate of each feature vector is obtained. Based on the probability density estimates of all feature vectors, the traffic density map is obtained.

[0032] Based on the kernel function and bandwidth parameters, density estimation is performed on each feature vector in the multimodal fusion features to obtain a probability density estimate for each feature vector. Based on all probability density estimates for all feature vectors, a traffic density map of the region to be predicted is obtained.

[0033] This invention calculates the traffic density map of multimodal fusion features based on kernel density estimation, which can effectively capture the local and global features of multimodal fusion features.

[0034] S400: Input the traffic volume density map and multimodal fusion features into the traffic volume prediction model to obtain the predicted traffic volume data output by the traffic volume prediction model.

[0035] The business volume prediction model is trained based on the sample business volume density map, sample multimodal fusion features, and the labels of the sample predicted business volume data, on the basis of the preset deep learning model.

[0036] The business volume prediction model is determined based on the following steps: obtaining labels for the sample predicted business volume data based on the sample business volume heatmap, the sample predicted business volume value, and the peak time period of the sample business volume; labeling the sample business volume density map and the sample multimodal fusion features based on the labels of the sample predicted business volume data to obtain labeled training samples; training, testing, and validating the preset deep learning model based on the training samples until the error output by the preset deep learning model is less than the set error, thus obtaining the business volume prediction model.

[0037] Based on existing historical traffic density data for the area to be predicted, a sample traffic density map is determined. Based on existing historical multimodal fusion characteristics for the area to be predicted, sample multimodal fusion characteristics are determined. Based on existing historical traffic heatmaps for the area to be predicted, a sample traffic heatmap is determined. Based on existing historical traffic values ​​for the area to be predicted, predicted sample traffic values ​​are determined. Based on existing historical peak traffic time periods for the area to be predicted, peak traffic time periods for the samples are determined.

[0038] Based on the sample traffic volume heatmap, predicted sample traffic volume values, and peak time periods for sample traffic, labels are obtained for the predicted sample traffic volume data. Based on these labels, the sample traffic volume density map and multimodal fusion features are labeled to obtain labeled training samples. The training samples are divided into training, testing, and validation sets according to a preset ratio. A preset deep learning model is constructed and trained using the training set. The trained deep learning model is tested using the testing set. The tested deep learning model is validated using the validation set. When the error output by the preset deep learning model is less than a set error, the model is considered validated, and the traffic volume prediction model is obtained.

[0039] The traffic volume density map and multimodal fusion features are input into the traffic volume prediction model. The traffic volume prediction model performs feature extraction and feature analysis on the traffic volume density map and multimodal fusion features to obtain the predicted traffic volume data for the region to be predicted in the future time period. Among them, feature analysis includes time series analysis.

[0040] Furthermore, the system displays predicted traffic volume data, including a traffic volume heatmap, predicted traffic volume values, and peak traffic periods. The network configuration is then updated based on this traffic volume data.

[0041] The network traffic prediction method provided by this invention obtains multimodal fusion features affecting network traffic based on multimodal data, achieving comprehensive collection of features influencing network traffic in the area to be predicted, which is beneficial to improving the accuracy of subsequent acquisition of traffic density maps and predicted traffic data. Kernel density estimation is performed on the multimodal fusion features based on kernel functions and bandwidth parameters to obtain a traffic density map, enabling accurate calculation of traffic density in the area to be predicted. Through a traffic prediction model, automatic feature extraction and analysis are performed on the traffic density map and multimodal fusion features to obtain predicted traffic data, improving the efficiency of obtaining predicted traffic data. This invention improves the accuracy of predicted traffic based on multimodal fusion features and kernel density estimation, and improves the efficiency of predicted traffic based on the traffic prediction model.

[0042] Based on the above embodiments, the multimodal data is acquired based on steps S110 to S140, and the specific steps are as follows.

[0043] S110: Based on satellite imagery and urban monitoring systems, acquire the geographical layout, building density, and traffic flow of the area to be predicted, and obtain video data of the area to be predicted.

[0044] S120: Based on the urban noise monitoring system, the environmental noise level of the area to be predicted is obtained at different times and locations, and the audio data of the area to be predicted is obtained.

[0045] S130: Based on publicly available data from public platforms, obtain text data that affects the communication needs of the area to be predicted.

[0046] S140: Multimodal data is obtained based on video data, audio data, and text data.

[0047] Video data of the area to be predicted is obtained by capturing information such as its geographical layout, building density, and traffic flow from satellite imagery and urban surveillance systems. This video data reflects the density and activity of network traffic in the area. Audio data of the area is obtained by acquiring environmental noise levels at different times and locations from urban noise monitoring systems. This audio data reflects the intensity of activity in the area (e.g., the city) through noise levels. Text data influencing communication needs in the area is obtained from publicly available data from public platforms, including social media platforms, news websites, and public organization platforms. This text data includes information on large-scale events, traffic accidents, or other events that may affect communication needs. Based on the video, audio, and text data, multimodal data is obtained.

[0048] Furthermore, multimodal data also includes historical traffic data, time stamps, meteorological information, and other multimodal data for the area to be predicted. For example, other multimodal data can be obtained from network operators and meteorological departments.

[0049] The multimodal data of this invention includes video data, audio data, and text data, which greatly expands the feature range for obtaining network traffic volume affecting the area to be predicted. It can comprehensively consider factors such as the geographical distribution, social activities, and environmental changes of the area to be predicted, which is conducive to improving the accuracy of obtaining traffic density maps and predicting traffic volume data.

[0050] Based on the above embodiments, feature extraction and feature fusion are performed on the multimodal data affecting network traffic in the region to be predicted to obtain multimodal fused features, including S150 to S190, and the specific steps are as follows.

[0051] S150: Preprocess the multimodal data to obtain preprocessed video data, preprocessed audio data, and preprocessed text data. Preprocessing includes uniform formatting, data cleaning, and standardization.

[0052] S160: An image feature processing module based on the contrastive language-image pre-trained CLIP model, which extracts features from pre-processed video data to obtain image spatial features.

[0053] S170: A text feature processing module based on the CLIP model, which extracts semantic and contextual features from preprocessed text data to obtain event semantic features.

[0054] S180: Based on the multilayer perceptron model, noise intensity features are extracted from the preprocessed audio data to obtain noise features.

[0055] S190: The image spatial features, event semantic features, and noise features are fused to obtain multimodal fusion features.

[0056] like Figure 2 As shown, all collected multimodal data undergoes unified formatting, data cleaning, and standardization to improve the quality of the multimodal data.

[0057] Contrastive Language-Image Pretraining (CLIP) is a deep learning model used to process both image and text information simultaneously. The main goal of CLIP is to embed images and text into the same feature space through contrastive learning, ensuring that similar images and text are close together, while unrelated images and text are far apart. The CLIP model consists of a visual encoder (image feature processing module) for processing images and a text encoder (text feature processing module) for processing text. These two encoders convert preprocessed video data and preprocessed text data into feature vectors (including image spatial features and event semantic features), respectively.

[0058] Noise features are obtained by extracting noise intensity features from the preprocessed audio data using a multi-layer perceptron (MLP) model.

[0059] Image spatial features, event semantic features, and noise features are input into the fusion layer of a neural network to perform data fusion, resulting in multimodal fusion features. Multimodal fusion features consist of multiple feature vectors.

[0060] This invention improves the accuracy of feature extraction from preprocessed video and text data by acquiring image spatial and event semantic features through the CLIP model. By fusing image spatial features, event semantic features, and noise features, multimodal fusion features are obtained, ensuring that features from all modalities are uniformly and effectively integrated, which is beneficial for improving the accuracy of subsequent acquisition of traffic density maps and prediction of traffic volume data.

[0061] Based on the above embodiments, the kernel function and bandwidth parameters are determined based on the distribution pattern of multimodal fusion features, including steps S210 to S230, each of which is as follows.

[0062] S210: When the distribution is normal, the kernel function is determined to be a Gaussian kernel function. The bandwidth parameter of the Gaussian kernel function is determined based on grid search and cross-validation.

[0063] S220: When the distribution is a concentrated distribution, the kernel function is determined to be the Epanechnikov kernel function. The bandwidth parameter of the Epanechnikov kernel function is determined based on grid search and cross-validation.

[0064] S230: When the distribution is a long-tailed distribution, the kernel function is determined to be the hyperbolic tangent kernel function. The bandwidth parameter of the hyperbolic tangent kernel function is determined based on grid search and cross-validation.

[0065] Based on grid search and cross-validation, the bandwidth parameter of the Gaussian kernel function is determined. Specifically, based on the numerical range of the multimodal fusion features, multiple initial bandwidth parameters are obtained. Each initial bandwidth parameter is used as a grid, and grid search and cross-validation are performed on each grid according to the multimodal fusion features and the Gaussian kernel function to obtain the evaluation value of each initial bandwidth parameter. The initial bandwidth parameter with the highest evaluation value is taken as the bandwidth parameter.

[0066] When the distribution is a normal distribution, the kernel function is a Gaussian kernel function, and the formula for calculating the Gaussian kernel function is as follows.

[0067] ; in, For kernel function, The feature vector is the feature vector of the multimodal fusion feature.

[0068] The multiple feature vectors of the multimodal fusion feature are divided into multiple subsets, and each subset is used as a validation set, while the remaining subsets are used as training sets.

[0069] Based on the numerical range of the multimodal fusion features, several initial bandwidth parameters of the Gaussian kernel function are obtained. The formula for calculating the initial bandwidth parameters of the Gaussian kernel function is as follows.

[0070] ; in, For the initial bandwidth parameters, This is the standard deviation of the training set (calculated based on the numerical range of the training set). This represents the number of samples in the training set.

[0071] Each initial bandwidth parameter is treated as a grid. Based on the validation set, training set, and Gaussian kernel function, grid search and cross-validation are performed on each grid to obtain the evaluation value of each initial bandwidth parameter. The initial bandwidth parameter with the highest evaluation value is taken as the bandwidth parameter.

[0072] Grid search and cross-validation include: estimating the kernel density based on the initial bandwidth parameters and Gaussian kernel function corresponding to the training set; calculating the evaluation value of the initial bandwidth parameters based on the validation set; transforming the training and validation sets; and calculating the evaluation value of each initial bandwidth parameter. The initial bandwidth parameter with the highest evaluation value is used as the bandwidth parameter of the Gaussian kernel function.

[0073] When the distribution is a concentrated distribution, the kernel function is the Epanechnikov kernel function, and the formula for calculating the Epanechnikov kernel function is as follows.

[0074] ; in, For kernel function, The feature vector is the feature vector of the multimodal fusion feature.

[0075] Based on the numerical range of the multimodal fusion features, multiple initial bandwidth parameters of the Epanechnikov kernel function are obtained. Each initial bandwidth parameter is treated as a grid, and grid search and cross-validation are performed on each grid according to the validation set, training set, and Epanechnikov kernel function to obtain an evaluation value for each initial bandwidth parameter. The initial bandwidth parameter with the highest evaluation value is then used as the bandwidth parameter of the Epanechnikov kernel function.

[0076] When the distribution has a long tail, the kernel function is the hyperbolic tangent kernel function. The formula for calculating the hyperbolic tangent kernel function is as follows.

[0077] ; in, For kernel function, The feature vector is the feature vector of the multimodal fusion feature.

[0078] Based on the numerical range of the multimodal fusion features, multiple initial bandwidth parameters of the hyperbolic tangent kernel function are obtained. Each initial bandwidth parameter is treated as a grid, and grid search and cross-validation are performed on each grid according to the validation set, training set, and hyperbolic tangent kernel function to obtain an evaluation value for each initial bandwidth parameter. The initial bandwidth parameter with the highest evaluation value is then used as the bandwidth parameter of the hyperbolic tangent kernel function.

[0079] This invention determines the kernel function and bandwidth parameters based on the distribution pattern and numerical range of multimodal fusion features, thereby achieving targeted kernel density estimation for multimodal fusion features with different distribution patterns and numerical ranges and improving the accuracy of kernel density estimation for multimodal fusion features.

[0080] The kernel density estimation method used in this invention can dynamically adjust the kernel function and bandwidth parameters. Combined with the real-time data update mechanism, it enables the business volume prediction model to quickly respond to changes in the environment and market, and provide business volume predictions in real time, which greatly improves the practical value and timeliness of the prediction.

[0081] This invention enables accurate and timely traffic forecasting, allowing wireless network operators to plan network resources, such as base station deployment and spectrum allocation, more scientifically, thereby reducing operating costs and improving resource utilization efficiency while ensuring service quality.

[0082] This invention is not only applicable to conventional wireless network traffic forecasting, but can also be extended to various scenarios such as emergency communication management and support for large-scale public events. For example, during large-scale events such as music festivals or sporting events, by accurately forecasting traffic volume in specific time periods and areas, network operators can adjust and optimize network configurations in advance to ensure the continuity and stability of communication services.

[0083] This invention significantly improves the end-user experience through more precise and timely network service adjustments. Especially during peak traffic periods and important public events, optimized network resource allocation can effectively prevent network congestion and improve service quality.

[0084] The detailed traffic volume analysis and forecasting results provided by this invention can offer strong data support to network planners and decision-makers, helping them make more scientific and rational decisions, especially in the face of rapidly changing market and technological environments.

[0085] The network traffic prediction device provided by the present invention is described below. The network traffic prediction device described below can be referred to in correspondence with the network traffic prediction method described above.

[0086] like Figure 3 As shown, a network traffic prediction device includes: a feature extraction module 301, used to extract and fuse features from multimodal data affecting network traffic in the region to be predicted, to obtain multimodal fused features.

[0087] The determination module 302 is used to determine the kernel function and bandwidth parameters based on the distribution pattern of multimodal fusion features.

[0088] The traffic density map acquisition module 303 is used to perform kernel density estimation on multimodal fusion features based on kernel function and bandwidth parameters to obtain a traffic density map.

[0089] The prediction module 304 is used to input the traffic volume density map and multimodal fusion features into the traffic volume prediction model and obtain the predicted traffic volume data output by the traffic volume prediction model. The traffic volume prediction model is trained based on the sample traffic volume density map, sample multimodal fusion features and the labels of the sample predicted traffic volume data on the basis of the preset deep learning model.

[0090] The network traffic prediction device provided by this invention acquires multimodal fusion features affecting network traffic based on multimodal data, achieving comprehensive collection of features influencing network traffic in the area to be predicted, which is beneficial to improving the accuracy of subsequent acquisition of traffic density maps and predicted traffic data. Kernel density estimation is performed on the multimodal fusion features based on kernel functions and bandwidth parameters to obtain a traffic density map, enabling accurate calculation of traffic density in the area to be predicted. Through a traffic prediction model, automatic feature extraction and analysis are performed on the traffic density map and multimodal fusion features to obtain predicted traffic data, improving the efficiency of obtaining predicted traffic data. This invention improves the accuracy of obtaining predicted traffic based on multimodal fusion features and kernel density estimation, and improves the efficiency of obtaining predicted traffic based on a traffic prediction model.

[0091] In one embodiment, the feature extraction module 301 is further configured to: obtain the geographical layout of the area to be predicted, the building density of the area to be predicted, and the traffic flow of the area to be predicted based on satellite imagery and urban monitoring systems, thereby obtaining video data of the area to be predicted; obtain the environmental noise levels of the area to be predicted at different times and locations based on urban noise monitoring systems, thereby obtaining audio data of the area to be predicted; obtain text data affecting the communication needs of the area to be predicted based on publicly available data from public platforms; and obtain multimodal data based on video data, audio data, and text data.

[0092] In one embodiment, the feature extraction module 301 is used to: preprocess the multimodal data to obtain preprocessed video data, preprocessed audio data, and preprocessed text data, wherein the preprocessing includes uniform formatting, data cleaning, and standardization; an image feature processing module based on a contrastive language-image pre-trained CLIP model to extract features from the preprocessed video data to obtain image spatial features; a text feature processing module based on the CLIP model to extract semantic and contextual features from the preprocessed text data to obtain event semantic features; to extract noise intensity features from the preprocessed audio data based on a multilayer perceptron model to obtain noise features; and to fuse the image spatial features, event semantic features, and noise features to obtain multimodal fusion features.

[0093] In one embodiment, the determining module 302 is configured to: determine the kernel function as a Gaussian kernel function when the distribution is normally distributed, and determine the bandwidth parameter of the Gaussian kernel function based on grid search and cross-validation; determine the kernel function as an Epanechnikov kernel function when the distribution is centered, and determine the bandwidth parameter of the Epanechnikov kernel function based on grid search and cross-validation; and determine the kernel function as a hyperbolic tangent kernel function when the distribution is long-tailed, and determine the bandwidth parameter of the hyperbolic tangent kernel function based on grid search and cross-validation.

[0094] In one embodiment, the multimodal fusion feature includes multiple feature vectors, and the traffic density map acquisition module 303 is used to: obtain the probability density estimate of each feature vector with the kernel function as the center and the bandwidth parameter as the radius; and obtain the traffic density map based on the probability density estimates of all feature vectors.

[0095] In one embodiment, the prediction module 304 is further configured to: obtain labels for the sample predicted business volume data based on the sample business volume heatmap, the sample business volume prediction value, and the sample business peak time period; label the sample business volume density map and the sample multimodal fusion features based on the labels of the sample predicted business volume data to obtain labeled training samples; and train, test, and verify a preset deep learning model based on the training samples until the error output by the preset deep learning model is less than a set error to obtain a business volume prediction model.

[0096] In one embodiment, the determining module 302 is used to: obtain multiple initial bandwidth parameters based on the numerical range of the multimodal fusion features; take each initial bandwidth parameter as a grid, perform grid search and cross-validation on each grid according to the multimodal fusion features and the Gaussian kernel function, and obtain the evaluation value of each initial bandwidth parameter; and take the initial bandwidth parameter with the highest evaluation value as the bandwidth parameter.

[0097] Figure 4 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 4As shown, the electronic device may include a processor 410, a communications interface 420, a memory 430, and a communication bus 440. The processor 410, communications interface 420, and memory 430 communicate with each other via the communication bus 440. The processor 410 can call logical instructions in the memory 430 to execute a network traffic prediction method. This method includes: extracting and fusing features from multimodal data affecting network traffic in the region to be predicted to obtain multimodal fused features; determining kernel functions and bandwidth parameters based on the distribution of the multimodal fused features; estimating kernel density of the multimodal fused features based on the kernel function and bandwidth parameters to obtain a traffic density map; and inputting the traffic density map and multimodal fused features into a traffic prediction model to obtain predicted traffic data output by the traffic prediction model. The traffic prediction model is trained based on the labels of sample traffic density maps, sample multimodal fused features, and sample predicted traffic data, using a pre-defined deep learning model.

[0098] Furthermore, the logical instructions in the aforementioned memory 430 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0099] In another aspect, the present invention also provides a non-transitory computer-readable storage medium storing a computer program thereon. When executed by a processor, the computer program implements a method for predicting network traffic volume provided by the methods described above. The method includes: extracting and fusing features from multimodal data affecting network traffic volume in the region to be predicted to obtain multimodal fused features; determining a kernel function and bandwidth parameters based on the distribution pattern of the multimodal fused features; estimating the kernel density of the multimodal fused features based on the kernel function and bandwidth parameters to obtain a traffic volume density map; and inputting the traffic volume density map and the multimodal fused features into a traffic volume prediction model to obtain predicted traffic volume data output by the traffic volume prediction model. The traffic volume prediction model is trained based on the labels of sample traffic volume density maps, sample multimodal fused features, and sample predicted traffic volume data, on the basis of a preset deep learning model.

[0100] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0101] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for predicting network traffic, characterized in that, include: Feature extraction and feature fusion are performed on multimodal data affecting network traffic in the region to be predicted to obtain multimodal fused features; the multimodal data includes video data, audio data, and text data; The obtained multimodal fusion features include: The image feature processing module based on the contrastive language-image pre-trained CLIP model extracts features from the pre-processed video data to obtain image spatial features; The text feature processing module based on the CLIP model extracts semantic and contextual features from the preprocessed text data to obtain event semantic features. The CLIP model is used to embed images and text into the same feature space through contrastive learning, so that similar images and texts are closer in the feature space, while unrelated images and texts are farther apart. Based on a multilayer perceptron model, noise intensity features are extracted from the preprocessed audio data to obtain noise features. The image spatial features, the event semantic features, and the noise features are fused to obtain the multimodal fusion features; Based on the distribution pattern of the multimodal fusion features, the kernel function and bandwidth parameters are determined; Based on the kernel function and the bandwidth parameter, the kernel density of the multimodal fusion features is estimated to obtain a traffic density map; The traffic volume density map and the multimodal fusion feature are input into the traffic volume prediction model to obtain the predicted traffic volume data output by the traffic volume prediction model. The traffic volume prediction model is trained based on the sample traffic volume density map, sample multimodal fusion feature and the label of the sample predicted traffic volume data, on the basis of a preset deep learning model.

2. The method for predicting network traffic according to claim 1, characterized in that, The multimodal data was obtained based on the following steps: Based on satellite imagery and urban monitoring systems, the geographical layout, building density, and traffic flow of the area to be predicted are obtained to generate video data of the area to be predicted. Based on the urban noise monitoring system, the environmental noise levels of the area to be predicted are obtained at different times and locations, and the audio data of the area to be predicted is obtained. Based on publicly available data from public platforms, obtain text data that influences the communication needs of the region to be predicted.

3. The method for predicting network traffic volume according to claim 1, characterized in that, The preprocessing includes uniform formatting, data cleaning, and standardization.

4. The method for predicting network traffic according to claim 1, characterized in that, The determination of the kernel function and bandwidth parameters based on the distribution pattern of the multimodal fusion features includes: When the distribution shape is a normal distribution, the kernel function is determined to be a Gaussian kernel function, and the bandwidth parameter of the Gaussian kernel function is determined based on grid search and cross-validation. When the distribution pattern is a concentrated distribution, the kernel function is determined to be the Epanechnikov kernel function, and the bandwidth parameter of the Epanechnikov kernel function is determined based on the grid search and the cross-validation. When the distribution pattern is a long-tailed distribution, the kernel function is determined to be a hyperbolic tangent kernel function. Based on the grid search and the cross-validation, the bandwidth parameter of the hyperbolic tangent kernel function is determined.

5. The method for predicting network traffic according to claim 1, characterized in that, The multimodal fusion feature includes multiple feature vectors. The step of estimating the kernel density of the multimodal fusion feature based on the kernel function and the bandwidth parameter to obtain a traffic density map includes: Using the kernel function as the center and the bandwidth parameter as the radius, obtain the probability density estimate of each feature vector; The traffic density map is obtained based on the probability density estimates of all the feature vectors.

6. The method for predicting network traffic according to claim 1, characterized in that, The business volume prediction model is determined based on the following steps: The labels for the sample predicted business volume data are obtained based on the sample business volume heatmap, the sample business volume prediction values, and the sample business peak time periods. Based on the labels of the sample predicted business volume data, the sample business volume density map and the sample multimodal fusion features are labeled to obtain labeled training samples. The preset deep learning model is trained, tested, and validated based on the training samples until the error output by the preset deep learning model is less than the set error, thus obtaining the business volume prediction model.

7. The method for predicting network traffic according to claim 4, characterized in that, The determination of the bandwidth parameter of the Gaussian kernel function based on grid search and cross-validation includes: Based on the numerical range of the multimodal fusion features, multiple initial bandwidth parameters are obtained; Using each of the initial bandwidth parameters as a grid, based on the multimodal fusion features and the Gaussian kernel function, perform grid search and cross-validation on each grid to obtain an evaluation value for each of the initial bandwidth parameters; The initial bandwidth parameter with the highest evaluation value is used as the bandwidth parameter.

8. A network traffic prediction device, characterized in that, include: The feature extraction module is used to extract and fuse features from multimodal data affecting network traffic in the region to be predicted, to obtain multimodal fused features; the multimodal data includes video data, audio data, and text data; The determination module is used to determine the kernel function and bandwidth parameters based on the distribution pattern of the multimodal fusion features; The traffic density map acquisition module is used to perform kernel density estimation on the multimodal fusion features based on the kernel function and the bandwidth parameter to obtain the traffic density map; The prediction module is used to input the traffic volume density map and the multimodal fusion feature into the traffic volume prediction model and obtain the predicted traffic volume data output by the traffic volume prediction model. The traffic volume prediction model is trained based on the sample traffic volume density map, sample multimodal fusion feature and the label of the sample predicted traffic volume data on the basis of a preset deep learning model. The feature extraction module is used to extract features from preprocessed video data based on the image feature processing module of the contrastive language-image pre-trained CLIP model to obtain image spatial features; the text feature processing module based on the CLIP model extracts semantic and contextual features from preprocessed text data to obtain event semantic features; the CLIP model is used to embed images and text into the same feature space through contrastive learning, so that similar images and texts are close in the feature space, while unrelated images and texts are far apart; noise intensity features are extracted from preprocessed audio data based on a multilayer perceptron model to obtain noise features; the image spatial features, the event semantic features, and the noise features are fused to obtain the multimodal fusion features.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the network traffic prediction method as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the network traffic prediction method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Flow-by-flow multi-mode network flow prediction method

    CN117155808A

  • Event occurrence time probability prediction method and system based on kernel density estimation

    CN118095565A