5G network self-optimization method and device based on artificial intelligence
Through artificial intelligence-based methods, multi-dimensional network data is acquired and processed in real time, and network prediction models and deep reinforcement learning optimization models are used to generate parameter adjustment instructions. This solves the problem that traditional 5G network optimization methods cannot respond to changes in network status in real time, and realizes adaptive optimization and performance improvement of the network.
Patent Information
- Application Number
- CN202511003500.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-21
- Publication Date
- 2025-09-05
AI Technical Summary
Traditional 5G network optimization methods cannot respond to network status changes in real time, resulting in degraded network performance and poor user experience.
Adopting an AI-based approach, multi-dimensional network data is acquired in real time, and data fusion and preprocessing are performed. Utilizing network prediction models and deep reinforcement learning optimization models, parameter adjustment instructions for network devices are generated to achieve adaptive optimization.
It realizes adaptive optimization of 5G networks, which can effectively respond to changes in network status and improve network performance and user experience.
Smart Images

Figure CN120602965A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a 5G network self-optimization method and device based on artificial intelligence. Background Art
[0002] 5G is the fifth generation of cellular mobile communication technology. Compared with 4G networks, 5G networks have higher bandwidth, lower latency, and more stable connections. They can support more users and more devices to access the network and provide higher quality services.
[0003] With the widespread adoption of 5G networks, their complexity and dynamism place higher demands on network optimization. Currently, traditional network optimization methods are typically based on static or semi-static strategies, setting fixed bandwidth allocation rules at the initial deployment stage and making only minor adjustments during subsequent operations. This approach offers the advantages of simplicity and manageability, but it lacks real-time response to network state changes, leading to degraded network performance and a poor user experience. Summary of the Invention
[0004] In view of this, the present invention provides a 5G network self-optimization method and device based on artificial intelligence, which can effectively respond to changes in network status and achieve adaptive optimization of the network.
[0005] A first aspect of the present invention provides a 5G network self-optimization method based on artificial intelligence, comprising:
[0006] Obtain multi-dimensional network data in real time;
[0007] fusing the multi-dimensional network data to obtain fused data;
[0008] Preprocessing the fused data to obtain time series features;
[0009] Input the time series features into the network prediction model, and output the predicted state of the network in the future;
[0010] Inputting the multi-dimensional network data and the predicted future state of the network into a deep reinforcement learning optimization model, and outputting an optimization result;
[0011] The optimization results are converted into actual parameter adjustment instructions for the network device.
[0012] Optionally, fusing the multi-dimensional network data to obtain fused data includes:
[0013] Aggregating all data in the multi-dimensional network data according to the same time granularity to obtain time slices;
[0014] determining a spatial anchor point in the aggregated data;
[0015] constructing an association key based on the time slice and the spatial anchor point;
[0016] Perform feature splicing based on the association key to obtain feature data;
[0017] Performing derivation based on the feature data to obtain derived feature data;
[0018] The feature data and the derived feature data are used as fused data.
[0019] Optionally, preprocessing the fused data to obtain time series features includes:
[0020] performing data cleaning on the fused data to obtain cleaned data;
[0021] Performing data conversion on the cleaned data to obtain standardized data;
[0022] Feature extraction is performed on the standardized data to obtain time series features.
[0023] Optionally, the network prediction model includes an input layer, a hidden layer, and an output layer. Inputting the time series features into the network prediction model and outputting the predicted future state of the network include:
[0024] The input layer flattens the time series features to obtain a feature vector;
[0025] The hidden layer performs full connection and nonlinear transformation of activation function on the feature vector to obtain hidden layer output;
[0026] The output layer determines the future predicted state of the network based on the output layer weights, the hidden layer outputs, and the bias term of the output layer.
[0027] Optionally, the deep reinforcement learning optimization model includes an online policy network, and the multi-dimensional network data and the predicted future state of the network are input into the deep reinforcement learning optimization model, and the optimization result is output, including:
[0028] Through the attention mechanism, the multi-dimensional network data and the predicted future state of the network and the historical action trajectory are integrated to obtain the state vector;
[0029] The state vector is input into the online policy network, and the action vector, i.e., the optimization result, is output.
[0030] Optionally, the deep reinforcement learning optimization model further includes an online value network, a target policy network, and a target value network. After inputting the state vector into the online policy network and outputting the optimization result, the model further includes:
[0031] Calculating the target action of the next state vector using the target strategy network;
[0032] Calculating the value corresponding to the target action using the target value network;
[0033] Calculate the mean square error based on the value corresponding to the target action to obtain the loss of the online value network;
[0034] Updating the online value network using the loss of the online value network to obtain an updated online value network;
[0035] The updated online value network is used to calculate the policy gradient, and the online policy network is updated based on the policy gradient to obtain an updated online policy network.
[0036] Optionally, converting the optimization result into an actual parameter adjustment instruction for the network device includes:
[0037] Analyzing the optimization results to obtain physical parameters;
[0038] According to the interface protocols of devices from different manufacturers, the physical parameters are adapted and processed to generate instructions that are compatible with the interface protocols.
[0039] A second aspect of the present invention provides a 5G network self-optimization device based on artificial intelligence, comprising:
[0040] An acquisition unit, used for acquiring multi-dimensional network data in real time;
[0041] A fusion unit, configured to fuse the multi-dimensional network data to obtain fused data;
[0042] A preprocessing unit, configured to preprocess the fused data to obtain time series features;
[0043] A prediction unit, configured to input the time series features into a network prediction model and output a predicted future state of the network;
[0044] an optimization unit, configured to input the multi-dimensional network data and the predicted future state of the network into a deep reinforcement learning optimization model, and output an optimization result;
[0045] The conversion unit is used to convert the optimization result into an actual parameter adjustment instruction of the network device.
[0046] Optionally, the fusion unit includes:
[0047] An aggregation unit, configured to aggregate all data in the multi-dimensional network data according to the same time granularity to obtain time slices;
[0048] a spatial anchor point determination unit, configured to determine a spatial anchor point in the aggregated data;
[0049] an association key construction unit, configured to construct an association key based on the time slice and the spatial anchor point;
[0050] A feature splicing unit, configured to perform feature splicing based on the association key to obtain feature data;
[0051] a derivation unit, configured to perform derivation based on the feature data to obtain derived feature data;
[0052] The fusion subunit is configured to use the feature data and the derived feature data as fused data.
[0053] Optionally, the pre-processing unit includes:
[0054] a data cleaning unit, configured to clean the fused data to obtain cleaned data;
[0055] A data conversion unit, configured to perform data conversion on the cleaned data to obtain standardized data;
[0056] The feature extraction unit is used to extract features from the standardized data to obtain time series features.
[0057] Optionally, the network prediction model includes an input layer, a hidden layer and an output layer;
[0058] The input layer is used to flatten the time series features to obtain feature vectors;
[0059] The hidden layer is used to perform a nonlinear transformation process of a full connection and an activation function on the feature vector to obtain a hidden layer output;
[0060] The output layer is used to determine the future prediction state of the network based on the output layer weights, the hidden layer outputs and the bias term of the output layer.
[0061] Optionally, the deep reinforcement learning optimization model includes an online policy network, and the optimization unit includes:
[0062] The state vector determination unit is used to obtain the state vector by fusing multi-dimensional network data, the network's future predicted state, and historical action trajectories through the attention mechanism;
[0063] The input unit is used to input the state vector into the online strategy network and output the action vector, i.e., the optimization result.
[0064] Optionally, the deep reinforcement learning optimization model further includes an online value network, a target strategy network, and a target value network, and the artificial intelligence-based 5G network self-optimization device further includes:
[0065] A first calculation unit, configured to calculate a target action for a next state vector using the target policy network;
[0066] A second calculation unit, configured to calculate the value corresponding to the target action using the target value network;
[0067] A third calculation unit is used to calculate the mean square error according to the value corresponding to the target action to obtain the loss of the online value network;
[0068] a first updating unit, configured to update the online value network using the loss of the online value network to obtain an updated online value network;
[0069] The second updating unit is used to calculate the policy gradient using the updated online value network, and update the online policy network based on the policy gradient to obtain an updated online policy network.
[0070] Optionally, the conversion unit includes:
[0071] An analysis unit, configured to analyze the optimization results to obtain physical parameters;
[0072] The adaptation unit is used to adapt the physical parameters according to the interface protocols of devices from different manufacturers and generate instructions that are compatible with the interface protocols.
[0073] A third aspect of the present invention provides an electronic device, comprising:
[0074] one or more processors;
[0075] a storage device having one or more programs stored thereon;
[0076] When the one or more programs are executed by the one or more processors, the one or more processors implement the artificial intelligence-based 5G network self-optimization method as described in any one of the first aspects.
[0077] A fourth aspect of the present invention provides a computer storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the artificial intelligence-based 5G network self-optimization method as described in any one of the first aspects is implemented.
[0078] As can be seen from the above scheme, the present invention provides an artificial intelligence-based 5G network self-optimization method and device. After acquiring multi-dimensional network data in real time, the multi-dimensional network data is fused and pre-processed to obtain time series features. The time series features are then analyzed using a network prediction model to obtain a predicted future state of the network. A deep reinforcement learning optimization model is then used to analyze the multi-dimensional network data and the predicted future state of the network to obtain an optimization result. Finally, the optimization result is converted into actual parameter adjustment instructions for network devices. This effectively responds to changes in network status and achieves the purpose of adaptive network optimization. BRIEF DESCRIPTION OF THE DRAWINGS
[0079] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.
[0080] Figure 1 A specific flow chart of an artificial intelligence-based 5G network self-optimization method provided in an embodiment of the present invention;
[0081] Figure 2 A flowchart of a data fusion method provided in another embodiment of the present invention;
[0082] Figure 3 A flowchart of a data preprocessing method provided by another embodiment of the present invention;
[0083] Figure 4 A schematic diagram of a deep reinforcement learning optimization model provided by another embodiment of the present invention;
[0084] Figure 5 A flowchart of a method for converting optimization results into adjustment instructions provided in another embodiment of the present invention;
[0085] Figure 6 A schematic diagram of an artificial intelligence-based 5G network self-optimization device provided in another embodiment of the present invention;
[0086] Figure 7 A schematic diagram of an electronic device for implementing an artificial intelligence-based 5G network self-optimization method provided in accordance with another embodiment of the present invention. DETAILED DESCRIPTION
[0087] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0088] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment," the term "another embodiment" means "at least one additional embodiment," and the term "some embodiments" means "at least some embodiments." Other terms are defined in the following description.
[0089] It should be noted that the information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the present invention are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions.
[0090] It should be noted that the concepts of "first" and "second" mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0091] It should be noted that the modifications of "one" and "multiple" mentioned in the present invention are illustrative rather than restrictive. Those skilled in the art should understand that unless otherwise clearly indicated in the context, it should be understood as "one or more".
[0092] The embodiment of the present invention provides a 5G network self-optimization method based on artificial intelligence, such as Figure 1 As shown, the specific steps include:
[0093] S101. Acquire multi-dimensional network data in real time.
[0094] Among them, multi-dimensional network data includes but is not limited to base station status, user behavior, network performance indicators, etc., which are not limited here.
[0095] In the specific implementation process of the present invention, multi-dimensional network data can be obtained in real time in base station equipment, user equipment, core network equipment, network management system, data acquisition equipment (such as probes, sensors, etc.), etc., which is not limited here.
[0096] Base stations are the core equipment of 5G networks, responsible for sending and receiving wireless signals. The following data can be collected in real time from base station equipment:
[0097] Base station status data, including but not limited to RSRP (reference signal received power), which is used to measure the signal strength received by the user equipment, SINR (signal to interference and noise ratio), which is used to measure signal quality and reflects the interference level, load information, which is used to measure the resource utilization of the base station (such as PRB utilization and number of users), and interference level, which is used to measure the strength of the interference signal received by the base station.
[0098] Network performance data, including but not limited to throughput, which is used to measure the data transmission rate of the base station; latency, which is used to measure the transmission delay of user data packets from the base station to the core network; and drop rate, which is used to measure the proportion of user connection interruptions.
[0099] User devices (such as mobile phones and IoT devices) are terminal devices of the 5G network and can collect the following data in real time:
[0100] User behavior data includes but is not limited to traffic data, i.e., the data usage of the user's device (such as uplink traffic and downlink traffic), movement trajectory data, i.e., the location information of the user's device (such as latitude and longitude, movement speed), and service type data, i.e., the type of service used by the user (such as video, game, voice).
[0101] Network performance data, including but not limited to RSRP, SINR, i.e., the signal strength and quality received by the user equipment, and handover success rate, i.e., the success rate of user equipment handover between different base stations.
[0102] Core network equipment is the control center of the 5G network, responsible for user authentication, data routing, and management. The following data can be collected in real time from core network equipment:
[0103] User behavior data, including but not limited to traffic data, i.e., the amount of data used by the user's device (such as uplink and downlink traffic), and service type data, i.e., the type of service used by the user (such as video, gaming, and voice).
[0104] Network performance data, including but not limited to latency, i.e., the transmission delay of user data packets from the base station to the core network; disconnection rate, i.e., the proportion of user connection interruptions; and QoS (quality of service), i.e., service quality indicators of user services (such as bandwidth, latency, and packet loss rate).
[0105] The network management system is responsible for monitoring and managing the operating status of the entire 5G network. The following data can be collected in real time from the network management system:
[0106] Base station status data, including but not limited to load data, i.e., base station resource utilization (such as PRB utilization and number of users), and interference level data, i.e., the strength of interference signals received by the base station.
[0107] Network performance data, including but not limited to throughput, i.e., the data transmission rate of the base station; latency, i.e., the transmission delay of user data packets from the base station to the core network; and drop rate, i.e., the proportion of user connection interruptions.
[0108] In 5G networks, specialized probes or sensor devices can also be deployed to collect specific data:
[0109] Network performance data includes but is not limited to latency, which is the transmission delay of user data packets measured by probes, and packet loss rate, which is the loss ratio of user data packets measured by probes.
[0110] User behavior data, including but not limited to traffic, i.e. measuring the data usage of user devices through probes, and service type data, i.e. identifying the service type used by users through probes.
[0111] S102: Fusing multi-dimensional network data to obtain fused data.
[0112] In the specific implementation process of the present invention, in order to achieve the integration of multi-dimensional data, it is necessary to fuse data from different devices, for example, to finally obtain B domain (business) data, O domain (operation) data and M domain (management) data.
[0113] B-domain (service) data primarily comes from user devices and core network equipment, reflecting user behavior and service usage. O-domain (operation) data primarily comes from base station equipment and network management systems, reflecting network operational status and performance. M-domain (management) data primarily comes from network management systems, reflecting network device configuration and management information.
[0114] In the specific application process of the present invention, an implementation of step S102 is as follows: Figure 2 Shown, including:
[0115] S201: Aggregate all data in the multi-dimensional network data according to the same time granularity to obtain time slices.
[0116] Specifically, all data is aggregated (averaged, summed, maximum, counted) or resampled to the same time granularity (e.g., 1 minute, 5 minutes, 15 minutes) to ensure that all features are within the same time slice.
[0117] S202: Determine a spatial anchor point in the aggregated data.
[0118] In the specific implementation process of the present invention, the following three spatial anchor points can be established but are not limited to:
[0119] Spatial Anchor 1: Cell. O-domain KPIs are naturally reported per cell. B-domain user locations (TAI / ECGI) can be mapped to serving cells. M-domain configurations are based on cells (cell parameters, home device).
[0120] Spatial Anchor 2: Network Device. Domain O alarms and device performance are device-based. Domain M topology and configuration are device-based. Domain B services flow through core network devices (which can be associated).
[0121] Spatial Anchor 3: User / Service Flow (Session / Flow). Domain B records service sessions (such as a video stream). Core Network / DPI provides the session path (base station -> transmission -> core network equipment), correlating O Domain KPIs with M Domain policies.
[0122] S203: Construct an association key based on the time slice and the spatial anchor point.
[0123] Specifically, a unique association key is created for each time slice. Common combinations are as follows:
[0124] [Timestamp, Cell_ID]: used for cell-level analysis.
[0125] [Timestamp, Device_ID]: used for device-level analysis.
[0126] [Timestamp, User_Hashed_ID]: used for user-level analysis (requires location mapping).
[0127] [Timestamp, Session_ID / Flow_ID]: used for business flow-level analysis (fine-grained, large data volume).
[0128] In the actual application process of the present invention, the association relationship may also be maintained, and the association relationship includes static association and dynamic association.
[0129] Static associations use static relationship tables to store core mapping relationships, for example:
[0130] Cell_Device_Map: cell ID->belonging base station device ID;
[0131] Device_Port_Topo: device ID->port ID->connected remote device / port;
[0132] Cell_Config: Cell ID->all M-domain configuration parameters;
[0133] User_Cell_History: The resident / serving cell of the user's Hashed ID in a time period (based on MR or location update).
[0134] Dynamic association can utilize, but is not limited to, core network / DPI session detail records (XDR), including: [Session_ID, Start_Time, End_Time, User_ID, App_Type, Source_IP, Destination_IP, Serving_Cell_ID, ...]. This can be associated with transmission paths and devices through IP or tunnel information, and signaling tracking (such as S1-U, N2 / N4) can be used to track the user / session path in the network. This is not limited here.
[0135] S204: Perform feature splicing based on the association key to obtain feature data.
[0136] Specifically, the B, O, and M field features under the same [Timestamp, Cell_ID] are horizontally spliced into one record.
[0137] Example fusion record (cell level, time granularity T): Timestamp, Cell_ID.
[0138] O domain characteristics: PRB_Util_UL, PRB_Util_DL, RRC_Conn_Users, Avg_UE_SINR, DL_Traffic, UL_Traffic, Handover_Success_Rate, Drop_Call_Rate, Active_Alarms(Count / Level);
[0139] B-domain characteristics (aggregated to the cell): Active_Users, Total_DL_Volume, Total_UL_Volume, Avg_Video_MOS, Video_Stutter_Rate, Avg_Web_Page_Load_Time, User_Distribution (e.g., Gold: 30%, Silver: 50%, Bronze: 20%);
[0140] M domain characteristics: Cell_Bandwidth, Carrier_Freq, Tx_Power, Antenna_Tilt, Neighbor_Cell_List, QCI_Config, Software_Version, Recent_Config_Change_Flag (eg, changein last 24h).
[0141] S205: Derivation is performed based on the feature data to obtain derived feature data.
[0142] Specifically, different domains have different derivation methods. For example, the O domain derives KPIs based on the mean, variance, and trend slope of the past N time windows; alarm correlation (overlapping of device / link alarms). The B domain derives services based on the proportion of services of different user tiers within the cell; and high-value user service experience indicators. The M domain derives services based on the difference between configuration parameters and standard / baseline values; and statistics on key parameters (such as the number of neighboring cells).
[0143] Of course, cross-domain interaction can also be derived, for example: service load vs. resource configuration: (DL_Traffic / Cell_Bandwidth), (Active_Users / Max_RRC_Conn); user experience vs. network performance: Avg_Video_MOS vs. Avg_UE_SINR vs. DL_Traffic.
[0144] S206: Use the feature data and the derived feature data as fused data.
[0145] S103: Preprocess the fused data to obtain time series features.
[0146] Among them, the preprocessing methods include but are not limited to data cleaning, conversion, feature extraction, etc., which are not limited here.
[0147] Optionally, in another embodiment of the present invention, an implementation of step S103 is as follows: Figure 3 Shown, including:
[0148] S301: Clean the fused data to obtain cleaned data.
[0149] Data cleaning is to deal with quality issues in the original data. It mainly includes data processing processes such as outlier detection and processing, missing value filling, noise filtering, duplicate data elimination, logical consistency verification, etc., and finally outputs a clean and complete original data set to ensure data quality.
[0150] Among them, the outlier identification method uses time series decomposition and isolation forest joint detection mechanism to identify abnormal data, including but not limited to seasonal decomposition of original time series data, separation of trend items, , Seasonal items and residual :
[0151] Where: X t is the original observation value at time point t (such as base station traffic), ω is the sliding window width (typical value ω=24, corresponding to a 24-hour period), is a moving average operator that extracts long-term trends, It is a seasonal decomposition algorithm.
[0152] The above decomposition is mainly used to separate the deterministic components in the data, so that the residual terms meet the independent and identically distributed assumptions, and provide a statistical basis for subsequent anomaly detection.
[0153] In the specific implementation process of the present invention, the isolation forest algorithm can also be applied to the residual term to calculate the anomaly score by constructing an isolation tree. The isolation forest algorithm is as follows:
[0154] ;
[0155] in, is the mean path length of sample x in the isolated tree, is the number of samples in the current window, is the path length normalization factor (0.5772 is Euler's constant).
[0156] when When it is judged as an anomaly, applying isolation forest on the residual term instead of the original data can avoid misjudgment caused by periodic fluctuations.
[0157] Among them, the missing value processing method implements the K nearest neighbor filling strategy based on time and space constraints; specifically includes:
[0158] (1) Construct a spatiotemporal correlation matrix to define the device topology neighborhood and time sliding window;
[0159] ;
[0160] in: For devices Geographical coordinates (latitude and longitude), is the spatial attenuation coefficient (typical value )、 is the time window size (typical value )、 is the indicator function (1 when the space-time constraints are satisfied).
[0161] The above processing can quantify the spatiotemporal correlation between devices, and devices with close distance and close time have higher weights.
[0162] (2) Search for K nearest neighbor samples within the spatiotemporal window;
[0163] (3) Generate filling value by weighted average:
[0164] ;
[0165] in: is the number of nearest neighbors (typical value =5)、X k For the The observations of the nearest neighbors, is the normalized weight derived from the spatiotemporal adjacency matrix A.
[0166] (4) When the missing rate exceeds the preset threshold, the data sample is directly deleted.
[0167] The noise suppression method implements adaptive low-pass filtering to reduce noise, retaining the main frequency signal (such as the flow baseline) and suppressing high-frequency noise (such as measurement jitter). The specific implementation method can be as follows:
[0168] (1) Design a Butterworth filter with adjustable cutoff frequency;
[0169] ;
[0170] Where: f is the signal frequency, is the cutoff frequency (dynamically adjusted), and N is the filter order (fixed value N=4).
[0171] (2) Dynamically adjust the cutoff frequency according to the signal spectrum characteristics:
[0172] ;
[0173] in: is the attenuation factor (empirical value =0.2)、 is the main frequency of the signal (estimated by power spectral density).
[0174] (3) Suppress high-frequency noise components and retain effective low-frequency signals.
[0175] ;
[0176] in: is the Fourier transform, is the inverse transform.
[0177] The logical consistency verification method establishes knowledge graph-driven verification rules, specifically including:
[0178] (1) Construct a network device topology knowledge graph and define entity relationship constraints; (2) Verify data logical consistency:
[0179] ;
[0180] in: 、 is the constraint function of relation r, .
[0181] (3) When data violates topological constraint rules, it is marked as inconsistent data.
[0182] The data deduplication method includes but is not limited to deploying a streaming data fingerprint deduplication mechanism, which is not limited here. The specific implementation methods include:
[0183] (1) Use SimHash algorithm to generate data fingerprint ;
[0184] ;
[0185] in: Represents the data feature vector, Represents feature weight vector, weight Dynamic allocation based on feature information entropy: ;
[0186] (2) Maintain the fingerprint Bloom filter within the sliding time window T_w;
[0187] ;
[0188] in: is an m-bit array, is a cluster of independent hash functions.
[0189] Window update mechanism: reset when Tw ends And launch a new window.
[0190] (3) When When present in the filter, it is considered duplicate data and discarded:
[0191] ;
[0192] in: For fingerprint The query results in the filter, It is the time difference between the current time and the time when the data arrives.
[0193] S302: Perform data conversion on the cleaned data to obtain standardized data.
[0194] Among them, data conversion is to convert data into a numerical form that can be processed by machines, which mainly includes but is not limited to classification data encoding, text vectorization, time series behavior sequence encoding processing, etc., and finally outputs a numerical / vectorized data set.
[0195] The specific implementation process of classification data encoding includes the following steps:
[0196] First, establish a network behavior classification system , the following conditions are met:
[0197] ;
[0198] in: The number of major categories of representative behaviors (e.g. : =Application interaction, =Media consumption, =Signaling control), Represents a sub-category (e.g. "Download APP", "Launch APP").
[0199] Then, each sub-category is coded to achieve the purpose of hierarchical coding mapping.
[0200] The specific implementation process of text vectorization includes the following steps:
[0201] First, build a communication field dictionary :
[0202] ;
[0203] in: Keywords (such as "5G", "QoS", "bandwidth"), calculate: ,in Represents the frequency of a word in a document; Represents the importance of a word, which is related to the frequency of occurrence of the word in the entire document collection.
[0204] Then, generate the embedding vector:
[0205] ;
[0206] in: Pre-trained embedding matrices for the domain, A collection of text segmentation.
[0207] The specific implementation process of temporal behavior sequence coding includes the following steps:
[0208] For each position i, we generate a vector of length d. Each dimension of this vector is calculated using a sine or cosine function. This allows the encodings of different positions to be correlated and periodic, thus introducing sequential order information into the model:
[0209] ;
[0210] in, is the position (starting from 0), indicating the position of the word in the sequence, is the behavior time step, representing the dimension index in the position encoding vector, is the encoding dimension, that is, the length of the encoding vector at each position.
[0211] The self-attention mechanism allows the model to focus on different parts of the sequence when processing the input sequence. Specifically, for each word in the input sequence, the model calculates the relationship between it and all other words, thereby weighting the information of all words. In this way, the model is able to capture contextual relationships:
[0212] ;
[0213] ;
[0214] in, is the matrix obtained by linear transformation of the input sequence, is the dimension of the key vector.
[0215] S303: Extract features from the standardized data to obtain time series features.
[0216] Among them, feature extraction is to extract high-order features from basic data and generate a feature matrix vector with high information density suitable for model input.
[0217] Specifically, it includes four steps: temporal feature extraction, spatial interaction feature calculation, graph structure feature generation, and feature cross-combination.
[0218] Among them, time series feature extraction can effectively capture local changes in time series, helping to better understand the dynamic characteristics of the data and thus improve the predictive ability of the model. Local statistical features are extracted from the data stream through sliding window statistics. Specifically, a fixed-size window is applied to the time series data, and certain statistics are calculated within the window. As the window slides, the dynamic characteristics of the data can be obtained:
[0219] ;
[0220] Here, ω is the window width (e.g., ω = 60 minutes), that is, the data range covered by each sliding window. The sliding step size (e.g. = 5 minutes) is the step size of the window sliding, which determines the time interval between each calculation.
[0221] Then, we generate high-order statistical features, such as mean, standard deviation, skewness, kurtosis, entropy, etc., to help reveal the nature of the data distribution. These statistics are very effective in describing the morphological characteristics of the data, especially when it is necessary to distinguish different types of data patterns:
[0222] ;
[0223] in, represents the mean, which is the mean of the data points in the window. ; It represents the standard deviation, which is an indicator to measure the volatility of data. The larger the standard deviation, the higher the degree of dispersion of the data. ; It stands for skewness, which describes the asymmetry of the data distribution. If the skewness is greater than 0, it means that the data distribution is biased to the right, otherwise it is biased to the left. ; It stands for excess kurtosis, which measures the “sharpness” of the data distribution. If the kurtosis is large, it means that the data distribution is concentrated; if it is small, it means that the data distribution is flat. ; stands for histogram entropy, which measures the degree of disorder of the data. .
[0224] in, is the probability of the b-th bucket; =10 means dividing the data into 10 buckets. A higher entropy value means the data is more uncertain or more evenly distributed.
[0225] The calculation of spatial interaction features can be implemented using, but is not limited to, the Geographical Decay Interaction Model, which calculates the interaction strength between devices based on spatial distance and time difference. This model takes into account two factors: spatial distance: The greater the physical distance between devices, the smaller the impact of the interaction. Time difference: The interaction between devices is affected by time; generally, the larger the time difference, the smaller the impact of the interaction.
[0226] ;
[0227] in: is the spatial distance between device i and device j (unit: km), which is usually the Euclidean distance or geographical distance between two points; is the time difference between device i and device j (unit: minutes), which indicates the time difference between the interaction between the two devices; α is the spatial attenuation coefficient, which is usually set to a constant (such as 1.5) to control the effect of distance on the interaction intensity. is the time attenuation coefficient, which is usually set to a constant (such as 0.02) to control the effect of time difference on interaction strength. is the key performance indicator (KPI) of device j, which can be certain performance data of the device, such as traffic, load, etc. is the set of neighbor devices of device i, that is, all devices that interact with device i.
[0228] In practical applications of this invention, graph structure feature generation can be achieved through, but not limited to, the multi-hop neighborhood aggregation method. This method aggregates node neighbor information to gradually build more complex node representations and can be applied to graph neural networks (GNNs). Through multi-hop aggregation, the model can capture both local and global structural information of nodes:
[0229] ;
[0230] in, is the representation of node j at step k. As the iteration proceeds, the node representation will contain more and more neighbor information. is the representation of node j in the k-1th step, that is, the node representation in the previous step, which is used to aggregate the neighbor information of the current node. is the set of neighbors of node i, that is, the set of nodes directly connected to node i. AGGREGATE is an aggregation function that can take various forms, including: Mean: takes the mean of the representations of neighboring nodes. MaxPooling: performs a max pooling operation on the representations of neighboring nodes. LSTM: Long Short-Term Memory (LSTM) network, used to process temporal information represented by neighboring nodes.
[0231] In the actual application of the present invention, a method of combining multiple features (such as time, space, graph structure, etc.) through tensor products may be used, but is not limited to, Tensor Outer Product Crossing.
[0232] Among them, the tensor product ( ) can be used to combine the multi-dimensional information of features to form a higher-dimensional representation by extending the traditional vector outer product method:
[0233] ;
[0234] in: They represent temporal features, spatial features and graph structure features respectively. Represents a tensor product operation, which combines the different dimensions of the three features into a new tensor. is the output tensor, with dimensions ,in: is the dimension of time feature, is the characteristic of the spatial dimension, is the dimension of the graph features.
[0235] In the actual application process of the present invention, it is also necessary to normalize the data to eliminate the dimension effect. For the O domain KPI sequence, a dynamic quantile normalization method can be used but is not limited to it, which is not limited here.
[0236] The dynamic quantile normalization method includes the following steps:
[0237] (1) Monitor the periodic mutation characteristics of network traffic data;
[0238] (2) Calculate dynamic quantiles based on the sliding time window, where is the quantile point;
[0239] (3) Perform normalization mapping:
[0240] ;
[0241] The technical meanings, data types, value ranges, and typical value examples of the parameter symbols in the above formulas can be found in Table 1 and are not limited here.
[0242] Table 1
[0243]
[0244] Among them, the dynamic quantile function: ; represents the data sample set within the time window t, is the inverse function of the empirical distribution function.
[0245] For the M-domain configuration parameters, a discrete parameter bucket normalization method may be used but is not limited to it, and is not limited here.
[0246] The discrete parameter bucket normalization method includes the following steps:
[0247] (1) Enumeration type configuration parameters (such as antenna tilt angle ) is divided into n discrete buckets;
[0248] (2) Establish a piecewise linear mapping function:
[0249] ;
[0250] The technical meanings, data types, value ranges, and typical value examples of the parameter symbols in the above formulas can be found in Table 2 and are not limited here.
[0251] Table 2
[0252]
[0253] Bucket boundary definition:
[0254] ;
[0255] when Time is included in bucket.
[0256] For the user behavior in domain B, a robust scaling method may be used but is not limited to it, which is not limited here.
[0257] The robust scaling method specifically includes the following steps:
[0258] (1) Calculate the interquartile range (IQR) of the user behavior indicator: Q_3 - Q_1;
[0259] (2) Perform anti-extreme value scaling:
[0260] ;
[0261] The technical meanings, data types, calculation methods, and example values of the parameter symbols in the above formulas can be found in Table 3 and are not limited here.
[0262] Table 3
[0263]
[0264] The outlier determination rules can be as follows:
[0265] ;
[0266] ;
[0267] When the above judgment conditions are met, Winsorize truncation processing can be used, which is not limited here.
[0268] For cross-domain fusion features, the graph embedding fusion normalization method can be used, but is not limited to, and specifically includes the following steps:
[0269] (1) Construct the cell-device topology graph G = (V, E);
[0270] (2) Generate entity embedding vectors through the Node2Vec algorithm ;
[0271] (3) Perform min-max normalization on the embedding vector:
[0272] ;
[0273] The technical meanings, data types, dimensions, and example values of the parameter symbols in the above formulas can be found in Table 4 and are not limited here.
[0274] Table 4
[0275]
[0276] Among them, the embedding generation algorithm: ; is the walking parameter, is the embedding dimension, and the extreme value is calculated as follows: .
[0277] S104: Input the time series features into the network prediction model, and output the predicted state of the network in the future.
[0278] Among them, the network prediction model includes an input layer, a hidden layer and an output layer.
[0279] Specifically, after the network prediction model receives the time series features, the input layer flattens the time series features to obtain a feature vector; the hidden layer performs a nonlinear transformation of the feature vector using full connection and activation functions to obtain the hidden layer output; the output layer determines the future prediction state of the network based on the output layer weights, the hidden layer output, and the output layer bias.
[0280] In the specific implementation process of the present invention, the time series features are fused time series features after LSTM or Transformer processing (different prediction models are used in different scenarios, LSTM processes the long dependencies of time series data; CNN is used for local feature extraction to help identify interference patterns; Transformer captures global features and long-distance dependencies to optimize service quality prediction). , where T is the number of time steps, is the feature dimension.
[0281] Among them, the input layer flattens the time series features to obtain the feature vector:
[0282] or (if already a vector), ;
[0283] For example, if , after flattening, .
[0284] Among them, each hidden layer undergoes nonlinear transformation through full connection and activation function:
[0285] ; is the weight matrix of the lth layer. is the bias vector. σ(⋅) is the activation function (such as ReLU, Sigmoid).
[0286] The output layer maps the hidden layer features to predicted values:
[0287] ;
[0288] in, is the output layer weight, which is used to map the output of the hidden layer to the output space. is the dimension of the output space, which depends on the type of task. For example, for classification tasks, is the number of categories; for regression tasks, is a scalar. is the feature dimension of the last hidden layer, which is usually consistent with the dimension of the model's internal representation space. L is the total number of hidden layers. It is the bias term of the output layer, ensuring that the model can have appropriate output when there is no input information. is the final prediction value of the model. Depending on the task, it can be a probability distribution for classification tasks or a numerical prediction for regression tasks.
[0289] During the model training process, a loss function is indispensable to evaluate the model performance. In the specific implementation process of the present invention, the mean squared error (MSE) can be used as the loss function.
[0290] The core of MSE is to find the difference between the model's prediction result and the actual target value, and to minimize MSE by adjusting the model's parameters so that its prediction value is as close to the actual target value as possible.
[0291] Suppose there is a dataset containing N samples. The output of the MLP model is the predicted value and the true value is . The calculation process of MSE is as follows:
[0292] Calculate the difference between the predicted value and the true value:
[0293] Error vector = ; Square each error value: ; take the average of all squared errors: .
[0294] In the specific implementation process of the present invention, the training method of the network prediction model can adopt, but is not limited to, supervised learning based on historical data and using gradient descent method to optimize model parameters. The steps are as follows:
[0295] Forward propagation: Input data passes through the model to obtain predicted values.
[0296] Calculate the loss: compare the MSE of the predicted value with the true value.
[0297] Backpropagation: Calculates the gradient of the loss with respect to the parameters. Care is taken to prevent gradient accumulation across batches, ensuring that gradients are calculated independently for each batch. The automatic differentiation system (Autograd) calculates the gradient of the loss with respect to each parameter. The chain rule propagates the error backward from the output layer, calculating gradients layer by layer.
[0298] Parameter update: The optimizer updates the weights based on the gradient.
[0299] Iteration: Repeat the above steps until convergence.
[0300] S105: Input the multi-dimensional network data and the predicted future state of the network into the deep reinforcement learning optimization model, and output the optimization result.
[0301] Specifically, the attention mechanism can be used to fuse multi-dimensional network data with the network's future predicted state and historical action trajectory to obtain a state vector. Then, the state vector is input into the online policy network, and the output is the action vector, which is the optimization result.
[0302] In the actual application of the present invention, the deep reinforcement learning optimization model is an intelligent agent constructed using a deep reinforcement learning algorithm (such as DQN, PPO, DDPG), which is designed with a state space, action space, reward function, model structure, and training method.
[0303] The state space defines the network environment information perceived by the agent, reflecting the current network state to support decision-making. It can realize multimodal state processing and feature fusion. Each state is a unique description of the environment. In 5G network optimization, the state may include information such as network load, user distribution, and signal strength. S consists of 3 parts. Real-time network status (Assuming there are K indicators), K indicators predicted by the network prediction model for the next H steps (a total of H*K dimensions), historical action trajectory A (assuming the last T actions, each action has A dimensions, then the historical action is T*A dimensions). Therefore, the total input dimension is: .
[0304] Among them, real-time network status The key performance indicators (KPIs) of the current network, such as latency, packet loss rate, and traffic, are recorded in the historical action trajectory A, which records previous action decisions, including adjustments to bandwidth, quality of service, and transmit power.
[0305] The state vector is obtained by fusing multi-source data through the attention mechanism :
[0306] ;
[0307] in: As input to the online policy network, Stored in the experience tuple ( )middle.
[0308] The action space of an agent is transformed or mapped to adapt to different environmental requirements or model architectures. It defines the set of actions that an agent can perform in a specific state, ensuring the feasibility of the actions and the support range of actual devices. In network optimization, actions may include adjusting base station power, resource allocation strategies, load balancing strategies, etc. It supports processing of discrete, linked, and hybrid actions.
[0309] Among them, base station power control includes: continuous action: adjusting the base station transmission power (such as continuous change within the range of ±2dBm) and discrete action: presetting power levels (such as low, medium and high).
[0310] Resource allocation strategies include: Spectrum allocation: Dynamically allocates the proportion of resource blocks in different frequency bands (such as millimeter wave or sub-6GHz). Time slot scheduling: Adjusts the TDD time slot ratio (such as the uplink / downlink time slot ratio).
[0311] Load balancing strategies include: User switching: triggering users to switch to adjacent base stations with lower loads (such as adjusting the handover threshold A3 offset). Traffic scheduling: migrating high-traffic users to frequency bands or cells with lighter loads.
[0312] The action space can include continuous actions as well as discrete actions.
[0313] Continuous actions can be expressed as: ,in, (Power adjustment amount). [0, 1] (proportion of millimeter wave resources). ∈[-6, +6]dB (switching threshold offset).
[0314] Discrete actions can be expressed as: ∈{Increase Power, Decrease Power, Allocate MoremmWave, Trigger Handover}, where, Strategic Network The output of OU noise is superimposed to generate the final action .
[0315] The reward function is used to define the value returned by the environment after the agent performs an action, which is used to evaluate the quality of the action. The reward function needs to quantify the network optimization goal and guide the agent to learn the optimal strategy. Multiple optimization goals need to be balanced to avoid conflicts. In network optimization, the reward function may be based on network performance indicators such as traffic, interference, load, etc. The core goals are throughput maximization (the reward is positively correlated with the total cell throughput), latency minimization (penalizing high latency), load balancing (penalizing load imbalance), and energy efficiency (penalizing high power consumption). The goal of this invention is multi-objective optimization, so a combined reward function containing three sub-rewards is defined:
[0316] 1. Performance Rewards: , used to penalize delay and packet loss rate.
[0317] in: The current latency (usually in milliseconds). The greater the latency, the worse the performance, so high latency needs to be penalized. is the packet loss rate at the current moment (usually expressed as a percentage). The higher the packet loss rate, the worse the network transmission quality, and the higher the packet loss rate needs to be penalized. α and β are weighting coefficients for latency and packet loss, respectively. They control the importance of latency and packet loss in the overall reward. The choice of α and β can be adjusted based on the application scenario. For example, if latency is more important than packet loss, α can be larger than β.
[0318] 2. Stable rewards: By penalizing the change between the current action and the previous action, we avoid frequent changes in network parameters and encourage smooth adjustment strategies. This helps reduce network instability and large fluctuations.
[0319] in, Indicates the network adjustment action at the current moment, for example, bandwidth or transmit power adjustment. Represents the network adjustment action at the previous moment.
[0320] 3. Prediction Rewards: This reward item encourages the model to make more accurate future state predictions by penalizing the error between the predicted value and the true value of the network prediction model.
[0321] in: is the actual network status at the current moment (such as the actual delay or packet loss rate). The network state at the future moment is predicted by the network prediction model. is a hyperparameter used to adjust the weight of the prediction reward. It controls how much the model focuses on prediction accuracy during optimization. A higher It will make the system pay more attention to prediction accuracy.
[0322] 4. The final comprehensive reward function ( ) is obtained by the weighted sum of performance reward, stability reward and prediction reward, balancing the contributions of the three sub-rewards to ensure the coordination of multi-objective optimization:
[0323] ;
[0324] Will be stored as the core element in the experience tuple ( ),in: is the weight of the performance reward, which is set to 0.6, meaning that the performance reward contributes the most to the final reward function. This indicates that the system prioritizes improving network performance. The weight of the reward for stability is set to 0.2, which means that stability also needs to be considered when optimizing network performance, but it is slightly lighter than performance. The weight of the predicted reward is set to 0.2, indicating that although the accuracy of the prediction is important, its contribution to the overall reward is relatively small. The system mainly focuses on performance and stability during optimization.
[0325] like Figure 4 As shown, in the specific implementation process of the present invention, the deep reinforcement learning optimization model can adopt the Actor-Critic framework, which includes four neural networks: an online policy network, an online value network, a target policy network, and a target value network.
[0326] The parameters and functions of the online strategy network, online value network, target strategy network, and target value network can be found in Table 5.
[0327] Table 5
[0328]
[0329] Online policy network (Actor), input is: state vector , dimension is d_s .
[0330] Online Strategy Network includes:
[0331] Fully connected layer 1: Input dimension d_s, output dimension 256, activation function ReLU. This layer uses the ReLU activation function to perform a nonlinear transformation on the input state vector. ReLU (Rectified Linear Unit) helps maintain sparsity and mitigate the vanishing gradient problem during network training by setting all negative values to zero and retaining positive values.
[0332] ;
[0333] in: is the weight matrix of the first layer. is the bias term of the first layer.
[0334] Fully connected layer 2: Input dimension 256 (from the previous layer), output dimension 256, activation function ReLU. This layer further transforms the hidden state to help the network learn more complex feature maps.
[0335] ;
[0336] in: is the weight matrix of the second layer. is the bias term of the second layer.
[0337] Fully connected layer 3: Input dimension 256, output dimension 256, activation function ReLU. This layer continues to extract features, further enhancing the network's expressive power and helping capture deeper state information.
[0338] ;
[0339] in: is the weight matrix of the third layer. is the bias term of the third layer.
[0340] Output layer: fully connected layer, input dimension 256 (output from the last layer), output action item a, dimension d_a.
[0341] ;
[0342] in: is the weight matrix of the fourth layer. is the bias term of the fourth layer.
[0343] Motion generation: The actual motion needs to be scaled according to the motion range. For example, if a motion component The actual range is [ , ], then the final action is: (The activation function tanh clamps each action component to [-1, 1]).
[0344] Input to the online value network (Critic): state vector s (dimension d_s) and action vector a (dimension d_a).
[0345] structure:
[0346] The state and action are concatenated into a vector [s, a] with a dimension of d_s + d_a. This concatenated vector contains the state information of the environment and the action information taken.
[0347] Fully connected layer 1: Input dimension d_s + d_a, output dimension 256, activation function ReLU. The first layer uses the ReLU activation function to perform a nonlinear mapping of the input state and action combination, helping the network extract useful features and enabling the network to learn more complex state-action associations.
[0348] ;
[0349] in: is the weight matrix of the first layer. is the bias term of the first layer.
[0350] Fully connected layer 2: Input dimension 256 (from the output of the previous layer), output dimension 256, activation function ReLU. Continues feature transformation to help the model more accurately understand the relationship between state and action.
[0351] ;
[0352] in: is the weight matrix of the second layer. is the bias term of the second layer.
[0353] Fully connected layer 3: Input dimension 256, output dimension 256, activation function ReLU. The third layer further enhances the representation capability of the network and further captures complex nonlinear patterns.
[0354] ;
[0355] in: is the weight matrix of the third layer. is the bias term of the third layer.
[0356] Output layer: A fully connected layer with an input dimension of 256 (the output from the last layer) and an output value (Q) of dimension 1 (no activation function). This Q-value represents the expected long-term cumulative reward for the current state and action. The lack of an activation function means that the range of output values is unrestricted, allowing the network to output any real-valued value representing the value of an action. This provides an effective way to evaluate the quality of the current action policy and adjust the policy accordingly.
[0357] ;
[0358] in: is the weight matrix of the fourth layer. is the bias term of the fourth layer.
[0359] Target network: Create target networks for Actor and Critic respectively. Their structures are the same as those of the online Actor and Critic networks, and their parameters are copied from the online network through soft updates.
[0360] During training, noise is added to the actions output by the Actor Network to explore the environment. Ornstein-Uhlenbeck (OU) noise or Gaussian noise is used. No noise is added during testing.
[0361] In the specific implementation process of the present invention, the training method of the deep reinforcement learning optimization model can adopt but is not limited to the following methods:
[0362] 1. Initialization: Orthogonally initialize the Actor / Critic network parameters.
[0363] 2. Collect experience:
[0364] (1) The agent selects actions through the policy network and exploration noise: ;
[0365] (2) Storing experience tuples ( ) to the replay pool.
[0366] 3. Random sampling:
[0367] Draw a small batch of samples from the replay pool ;
[0368] 4. Update the value network:
[0369] (1) Use the target policy network to calculate the target action for the next state vector:
[0370] ;
[0371] (2) Use the target value network to calculate the Q value corresponding to the target action:
[0372] ;
[0373] in: is the immediate reward after taking action i at the current moment. γ is a discount factor, which indicates the influence of future rewards. It is usually between 0 ≤ γ ≤ 1. The smaller γ is, the less influence future rewards have on the current decision. is calculated by the target value network, the next state and target action The Q value under .
[0374] (3) Use mean square error (MSE) to optimize the loss of the online value network:
[0375] ;
[0376] (4) Update the online value network using the loss of the online value network to obtain the updated online value network:
[0377] ;
[0378] in: The learning rate controls the step size of parameter updates and determines the magnitude of parameter adjustments during each update. A larger learning rate may lead to unstable training, while a smaller learning rate may slow convergence. is the loss function Relative to the main value network parameters The gradient of . It reflects the direction of change of the loss function relative to the network parameters and is used to guide parameter updates.
[0379] 5. Use the updated online value network to calculate the policy gradient, and update the online policy network based on the policy gradient to obtain the updated online policy network:
[0380] Use the online value network to calculate the policy gradient and maximize expected value.
[0381] Gradient Ascent: ;in: .
[0382] It can be seen that each step of the model training involves the main network parameters ( ) is updated. In the actual application process of the present invention, the target value network and the target strategy network will also be updated. This can be done slowly through but not limited to soft updates to ensure smooth training.
[0383] 1. Soft update of the target network (target value network, target policy network): Each time the parameters of the main network (online value network, online policy network) are updated, the target network will be soft updated at a smaller ratio.
[0384] ;
[0385] ;
[0386] in: This is a hyperparameter for soft updates, typically set to a small value (e.g., 0.001). It controls the update step size between the target network and the main network. A small τ value helps smooth the transition and avoids drastic changes in the target network during training.
[0387] Time-varying update coefficient :
[0388] ;
[0389] in: is the initial update coefficient, that is, the initial target network update step size, usually a small positive value. is the attenuation factor used to control The rate of decay, , which determines how fast the coefficient decreases, and a larger value will lead to faster decay.
[0390] 2. Value Network : Adjust the network by minimizing the error between the target Q value and the current Q value.
[0391] 3. Policy Network : Adjust the strategy by maximizing the Q value evaluated by the value network.
[0392] S106: Convert the optimization result into actual parameter adjustment instructions for the network device.
[0393] Specifically, the optimization results output by the deep reinforcement learning model are converted into parameter adjustment instructions for actual network equipment and applied to the 5G network, such as adjusting base station power, resource allocation strategy, load balancing strategy, etc., and ensuring that the adjustment process is efficient, safe, and reliable.
[0394] Optionally, in another embodiment of the present invention, an implementation of step S106 is as follows: Figure 5 Shown, including:
[0395] S501: Analyze the optimization results to obtain physical parameters.
[0396] In the specific implementation process of the present invention, parsing the optimization results mainly includes the following steps: identifying the action type in the optimization results, parsing the action vector, mapping the parsed action, and finally obtaining specific physical parameters.
[0397] Action types include continuous, discrete, and mixed. The action space for continuous actions is a continuous space of values. For example, base station power adjustment. The action space for discrete actions is discrete, meaning the agent can only choose an action from a limited number of options. For example, the action space might include "close," "open," and so on. In the action space of mixed actions, the system contains both discrete and continuous actions.
[0398] Methods for parsing action vectors include, but are not limited to, decomposing the action vector: decomposing the original action vector into different parts, corresponding to discrete or continuous actions. For mixed action spaces, the parser needs to distinguish between discrete and continuous parts and process them differently depending on their type. Mapping to the action space of the deep reinforcement learning optimization model: Converting the parsed vector into a specific operation through a mapping relationship with the defined action space. For continuous actions, this may involve scaling or normalizing the vector; for discrete actions, it may be an index lookup process.
[0399] The parsed actions are mapped, including discrete action space mapping, continuous action space mapping, and mixed action space mapping.
[0400] Spectrum allocation for discrete actions: The output of the deep reinforcement learning optimization model can be directly mapped to which spectrum resource to choose. Each action represents the selection of a specific spectrum block for allocation.
[0401] Base station access for discrete actions: The output of a deep reinforcement learning optimization model can be mapped to an access decision between a user and a base station. For example, if there are multiple base stations, the discrete actions can be mapped to the selection of each base station.
[0402] Load balancing for discrete actions: Discrete actions can selectively transfer traffic from one base station to another to optimize load distribution.
[0403] Power control for continuous actions: A deep reinforcement learning optimization model can output a power value (e.g., power intensity) that can be directly used to adjust the power output of a base station or device to optimize coverage or signal quality.
[0404] Channel selection for continuous actions: A deep reinforcement learning optimization model can output a frequency or time slot value indicating which specific channel to choose for communication, thereby optimizing signal quality and interference management.
[0405] Traffic regulation for continuous actions: The output continuous value can be mapped to the traffic allocation ratio for specific users or base stations, optimizing the utilization of network bandwidth.
[0406] In some cases, the output of a deep reinforcement learning optimization model may be a mixed action space, involving both discrete and continuous decisions. In this case, the action parser needs to combine discrete actions (such as base station and channel selection) with continuous actions (such as power control) for processing. For example:
[0407] The deep reinforcement learning optimization model outputs a discrete value representing the selected base station or channel, and a continuous value representing the corresponding power or bandwidth allocation. The action parser combines these two components to form the final network optimization decision.
[0408] S502: Adapt the physical parameters according to the interface protocols of devices from different manufacturers, and generate instructions that are compatible with the interface protocols.
[0409] It is understandable that the interface protocols of devices from different manufacturers may be different, so it is necessary to adapt the physical parameters and generate instructions that are compatible with the interface protocol.
[0410] For example: standardized interface:
[0411] O-RAN Alliance: Uses standardized interfaces (such as E2 and A1) of the Open Radio Access Network (O-RAN) to support unified control of multi-vendor equipment.
[0412] NETCONF / YANG model: Configure device parameters using the NETCONF protocol and the YANG data model.
[0413] Vendor-specific APIs:
[0414] For traditional devices that do not support O-RAN, call the API provided by the manufacturer (such as Huawei's iManager M2000 and Ericsson's ENM).
[0415] In the specific implementation process of the present invention, in order to avoid network shock or service interruption, it is also necessary to ensure that the timing and method of action execution are reasonable.
[0416] Execution can be done in real time, with latency-sensitive parameters (such as power adjustments) taking effect immediately. It can also be scheduled, with high-impact operations (such as frequency switching) executed during low-load periods. Batch execution can also be used, with large-scale adjustments executed in batches (such as adjusting 10% of base stations at a time) to avoid instability caused by simultaneous changes across the entire network. There are no restrictions here.
[0417] In the specific implementation process of the present invention, a rollback mechanism may also be included, for example, monitoring key indicators after execution (such as offline rate), and automatically rolling back to the previous configuration if a threshold is exceeded.
[0418] As can be seen from the above scheme, the present invention provides an artificial intelligence-based 5G network self-optimization method. After acquiring multi-dimensional network data in real time, the multi-dimensional network data is fused and pre-processed to obtain time series features. The time series features are then analyzed using a network prediction model to obtain a predicted future state of the network. A deep reinforcement learning optimization model is then used to analyze the multi-dimensional network data and the predicted future state of the network to obtain an optimization result. Finally, the optimization result is converted into actual parameter adjustment instructions for network devices. This effectively responds to changes in network status and achieves the goal of adaptive network optimization.
[0419] Another embodiment of the present invention provides a 5G network self-optimization device based on artificial intelligence, such as Figure 6 As shown, specifically including:
[0420] The acquisition unit 601 is used to acquire multi-dimensional network data in real time.
[0421] The fusion unit 602 is configured to fuse the multi-dimensional network data to obtain fused data.
[0422] Optionally, in another embodiment of the present invention, an implementation of the fusion unit 602 includes:
[0423] The aggregation unit is used to aggregate all data in the multi-dimensional network data according to the same time granularity to obtain time slices.
[0424] The spatial anchor point determination unit is used to determine the spatial anchor points in the aggregated data.
[0425] The association key construction unit is used to construct association keys based on time slices and spatial anchor points.
[0426] The feature splicing unit is used to perform feature splicing based on the association key to obtain feature data.
[0427] The derivation unit is used to perform derivation based on the feature data to obtain derived feature data.
[0428] The fusion subunit is used to use the feature data and the derived feature data as fused data.
[0429] The specific working process of the units disclosed in the above embodiments of the present invention can be found in the corresponding method embodiments, such as Figure 2 As shown, no further details are given here.
[0430] The preprocessing unit 603 is used to preprocess the fused data to obtain time series features.
[0431] Optionally, in another embodiment of the present invention, an implementation of the pre-processing unit 603 includes:
[0432] The data cleaning unit is used to clean the fused data to obtain cleaned data.
[0433] The data conversion unit is used to convert the cleaned data to obtain standardized data.
[0434] The feature extraction unit is used to extract features from the standardized data to obtain time series features.
[0435] The specific working process of the units disclosed in the above embodiments of the present invention can be found in the corresponding method embodiments, such as Figure 3 As shown, no further details are given here.
[0436] The prediction unit 604 is used to input the time series features into the network prediction model and output the predicted state of the network in the future.
[0437] Optionally, in another embodiment of the present invention, the network prediction model includes an input layer, a hidden layer, and an output layer;
[0438] The input layer is used to flatten the time series features to obtain feature vectors.
[0439] The hidden layer is used to perform full connection and nonlinear transformation of the feature vector using the activation function to obtain the hidden layer output.
[0440] The output layer is used to determine the future predicted state of the network based on the output layer weights, hidden layer outputs, and the output layer bias.
[0441] The specific working process of the units disclosed in the above embodiments of the present invention can be found in the corresponding method embodiments, which will not be repeated here.
[0442] The optimization unit 605 is used to input the multi-dimensional network data and the predicted future state of the network into the deep reinforcement learning optimization model and output the optimization result.
[0443] Optionally, in another embodiment of the present invention, the deep reinforcement learning optimization model includes an online policy network, and an implementation of the optimization unit 605 includes:
[0444] The state vector determination unit is used to obtain the state vector by fusing multi-dimensional network data, the network's future predicted state, and historical action trajectories through the attention mechanism.
[0445] The input unit is used to input the state vector into the online policy network and output the action vector, which is the optimization result.
[0446] The specific working process of the units disclosed in the above embodiments of the present invention can be found in the corresponding method embodiments, which will not be repeated here.
[0447] Optionally, in another embodiment of the present invention, the deep reinforcement learning optimization model further includes an online value network, a target policy network, and a target value network. An implementation of the artificial intelligence-based 5G network self-optimization device includes:
[0448] The first calculation unit is used to calculate the target action of the next state vector using the target policy network.
[0449] The second calculation unit is used to calculate the value corresponding to the target action using the target value network.
[0450] The third calculation unit is used to calculate the mean square error according to the value corresponding to the target action to obtain the loss of the online value network.
[0451] The first updating unit is configured to update the online value network using the loss of the online value network to obtain an updated online value network.
[0452] The second updating unit is used to calculate the policy gradient using the updated online value network, and update the online policy network based on the policy gradient to obtain an updated online policy network.
[0453] The specific working process of the units disclosed in the above embodiments of the present invention can be found in the corresponding method embodiments, which will not be repeated here.
[0454] The conversion unit 606 is configured to convert the optimization result into an actual parameter adjustment instruction for the network device.
[0455] The specific working process of the units disclosed in the above embodiments of the present invention can be found in the corresponding method embodiments, such as Figure 1 As shown, no further details are given here.
[0456] Optionally, in another embodiment of the present invention, an implementation of the conversion unit 606 includes:
[0457] The analytical unit is used to analyze the optimization results and obtain physical parameters.
[0458] The adaptation unit is used to adapt the physical parameters according to the interface protocols of devices from different manufacturers and generate instructions that are compatible with the interface protocols.
[0459] The specific working process of the units disclosed in the above embodiments of the present invention can be found in the corresponding method embodiments, such as Figure 5 As shown, no further details are given here.
[0460] As can be seen from the above scheme, the present invention provides an artificial intelligence-based 5G network self-optimization device. After acquiring multi-dimensional network data in real time, it fuses and pre-processes the multi-dimensional network data to obtain time series features. The time series features are then analyzed using a network prediction model to obtain a predicted future state of the network. A deep reinforcement learning optimization model is then used to analyze the multi-dimensional network data and the predicted future state of the network to obtain an optimization result. Finally, the optimization result is converted into actual parameter adjustment instructions for network devices. This effectively responds to changes in network status and achieves the goal of adaptive network optimization.
[0461] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), and the like.
[0462] Another embodiment of the present invention provides an electronic device, such as Figure 7 Shown, including:
[0463] One or more processors 701.
[0464] The storage device 702 stores one or more programs.
[0465] When the one or more programs are executed by the one or more processors 701, the one or more processors 701 implement the artificial intelligence-based 5G network self-optimization method as described in the above embodiments.
[0466] Another embodiment of the present invention provides a computer storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the artificial intelligence-based 5G network self-optimization method as described in the above embodiment is implemented.
[0467] In the context of the present invention, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0468] It should be noted that the computer-readable medium described above in the present invention may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. Computer-readable storage media may include, but are not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present invention, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present invention, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. This propagated data signal may take a variety of forms, including, but not limited to, electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wire, optical cable, RF (radio frequency), or any suitable combination thereof.
[0469] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.
[0470] Another embodiment of the present invention provides a computer program product, which, when executed, is used to perform the above-mentioned artificial intelligence-based 5G network self-optimization method.
[0471] In particular, according to an embodiment of the present invention, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processing device, the above-mentioned functions defined in the method of the embodiment of the present invention are performed.
[0472] Although the subject matter has been described in terms of structural features and / or method logic actions, it should be understood that the subject matter defined in the present invention is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the present invention.
[0473] Although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present invention. Certain features described in the context of a separate embodiment may also be implemented in combination in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented in multiple embodiments individually or in any suitable sub-combination.
[0474] The above description is merely a preferred embodiment of the present invention and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of application of the present invention is not limited to technical solutions formed by specific combinations of the aforementioned technical features. It also encompasses other technical solutions formed by any combination of the aforementioned technical features or their equivalents, without departing from the aforementioned application concepts. For example, a technical solution formed by replacing the aforementioned features with (but not limited to) technical features with similar functions in the present invention.
Claims
1. A 5G network self-optimization method based on artificial intelligence, characterized in that: include: Obtain multi-dimensional network data in real time; fusing the multi-dimensional network data to obtain fused data; Preprocessing the fused data to obtain time series features; Input the time series features into the network prediction model, and output the predicted state of the network in the future; Inputting the multi-dimensional network data and the predicted future state of the network into a deep reinforcement learning optimization model, and outputting an optimization result; The optimization results are converted into actual parameter adjustment instructions for the network device.
2. The 5G network self-optimization method based on artificial intelligence according to claim 1 is characterized in that: The fusing the multi-dimensional network data to obtain fused data includes: Aggregating all data in the multi-dimensional network data according to the same time granularity to obtain time slices; determining a spatial anchor point in the aggregated data; constructing an association key based on the time slice and the spatial anchor point; Perform feature splicing based on the association key to obtain feature data; Performing derivation based on the feature data to obtain derived feature data; The feature data and the derived feature data are used as fused data.
3. The 5G network self-optimization method based on artificial intelligence according to claim 1, characterized in that The preprocessing of the fused data to obtain time series features includes: performing data cleaning on the fused data to obtain cleaned data; Performing data conversion on the cleaned data to obtain standardized data; Feature extraction is performed on the standardized data to obtain time series features.
4. The 5G network self-optimization method based on artificial intelligence according to claim 1, characterized in that The network prediction model includes an input layer, a hidden layer, and an output layer. The time series features are input into the network prediction model, and the future prediction state of the network is obtained as an output, including: The input layer flattens the time series features to obtain a feature vector; The hidden layer performs full connection and nonlinear transformation of activation function on the feature vector to obtain hidden layer output; The output layer determines the future predicted state of the network based on the output layer weights, the hidden layer outputs, and the bias term of the output layer.
5. The 5G network self-optimization method based on artificial intelligence according to claim 1, characterized in that The deep reinforcement learning optimization model includes an online policy network, and the multi-dimensional network data and the predicted future state of the network are input into the deep reinforcement learning optimization model, and the optimization result is output, including: Through the attention mechanism, the multi-dimensional network data and the predicted future state of the network and the historical action trajectory are integrated to obtain the state vector; The state vector is input into the online policy network, and the action vector, i.e., the optimization result, is output.
6. The 5G network self-optimization method based on artificial intelligence according to claim 5, characterized in that The deep reinforcement learning optimization model further includes an online value network, a target policy network, and a target value network. After inputting the state vector into the online policy network and outputting the optimization result, the model further includes: Calculating the target action of the next state vector using the target strategy network; Calculating the value corresponding to the target action using the target value network; Calculate the mean square error based on the value corresponding to the target action to obtain the loss of the online value network; Updating the online value network using the loss of the online value network to obtain an updated online value network; The updated online value network is used to calculate the policy gradient, and the online policy network is updated based on the policy gradient to obtain an updated online policy network.
7. The 5G network self-optimization method based on artificial intelligence according to claim 1, characterized in that: Converting the optimization result into an actual parameter adjustment instruction for the network device includes: Analyzing the optimization results to obtain physical parameters; According to the interface protocols of devices from different manufacturers, the physical parameters are adapted and processed to generate instructions that are compatible with the interface protocols.
8. A 5G network self-optimization device based on artificial intelligence, characterized in that: include: An acquisition unit, used for acquiring multi-dimensional network data in real time; A fusion unit, configured to fuse the multi-dimensional network data to obtain fused data; A preprocessing unit, configured to preprocess the fused data to obtain time series features; A prediction unit, configured to input the time series features into a network prediction model and output a predicted future state of the network; an optimization unit, configured to input the multi-dimensional network data and the predicted future state of the network into a deep reinforcement learning optimization model, and output an optimization result; The conversion unit is used to convert the optimization result into an actual parameter adjustment instruction of the network device.
9. The 5G network self-optimization device based on artificial intelligence according to claim 8, characterized in that The fusion unit comprises: An aggregation unit, configured to aggregate all data in the multi-dimensional network data according to the same time granularity to obtain time slices; a spatial anchor point determination unit, configured to determine a spatial anchor point in the aggregated data; an association key construction unit, configured to construct an association key based on the time slice and the spatial anchor point; A feature splicing unit, configured to perform feature splicing based on the association key to obtain feature data; a derivation unit, configured to perform derivation based on the feature data to obtain derived feature data; The fusion subunit is configured to use the feature data and the derived feature data as fused data.
10. The 5G network self-optimization device based on artificial intelligence according to claim 8, characterized in that The pre-processing unit comprises: a data cleaning unit, configured to clean the fused data to obtain cleaned data; A data conversion unit, configured to perform data conversion on the cleaned data to obtain standardized data; The feature extraction unit is used to extract features from the standardized data to obtain time series features.
Citation Information
Patent Citations
Virtual resource dynamic capacity expansion and contraction method based on flow prediction and deep reinforcement learning
CN113810954A
Wireless routing optimization method based on attention mechanism and deep reinforcement learning
CN114423061A
Concentrator optimal configuration method and system based on deep learning
CN118659972A
Intranet service quality optimization method and system based on deep reinforcement learning
CN119496716A
Network traffic scheduling optimization method based on deep learning
CN120281665A
Cited By
5G network dynamic security capability scheduling method based on deep learning
CN121078436A