Transfer learning of network traffic prediction models between cellular base stations
Through transfer learning technology, using similarity networks and importance score matrices, the inaccuracy problem of base station service prediction in cellular communication networks is solved, the efficient allocation of base station spectrum is achieved, user needs are met, and the utilization efficiency of data and bandwidth is improved.
Patent Information
- Application Number
- CN202180055885.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2021-08-02
- Filing Date
- 2021-08-13
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2041-08-13
AI Technical Summary
In cellular communication networks, existing technologies have difficulty effectively utilizing historical data to perform base station service forecasts, resulting in improper or inefficient spectrum allocation that fails to meet the needs of user devices.
Through transfer learning technology, using similarity networks and importance score matrices, the prediction model of the target base station is trained based on the historical data of the source base station. The training process is adjusted to take the importance of parameters into account. Base stations with rich data are selected as source base stations for model migration and fine-tuning.
It improves the accuracy of base station spectrum allocation, reduces prediction errors, ensures that base stations can allocate resources in a timely and appropriate manner to meet user needs, and improves data and bandwidth efficiency.
Smart Images

Figure CN116097704B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the prediction of future communication traffic so that base stations can be appropriately configured. Background Art
[0002] This application relates to supporting edge operations for cellular communications. Traffic volume (or simply traffic), one of the most fundamental metrics of cellular communication networks, is a key reference variable. An example of cellular communication is 5G. Future / forecasted traffic is already being used to guide 5G operations, such as predictive resource allocation, dynamic spectrum management, and automated network slicing. Summary of the Invention
[0003] Solution to the problem
[0004] This disclosure relates to transfer learning based on predictions of similarities between a source base station and a target base station. Parameter importance is determined and training is adjusted to account for this. Lack of historical data is compensated for by selecting a base station with extensive historical data as the source base station.
[0005] Provided herein is a server configured to manage traffic prediction model transfer learning between cellular communication base stations (including, as a non-limiting example, 5G base stations), the server comprising: one or more processors; and one or more memories storing a program, wherein execution of the program by the one or more processors is configured to cause the server to at least: receive a first plurality of base station statistics, wherein the first plurality of base station statistics comprises a first data set of a first size from a first base station; receive a second plurality of base station statistics, wherein the second plurality of base station statistics comprises a second data set of a second size corresponding to a second base station; select the first base station as a source base station; train a similarity network; receive a source prediction model and a first importance score matrix from the first base station; receive a prediction model request from a target base station, wherein the target base station is the second base station; calculate a first similarity using the similarity network; obtain a first scaled importance score moment based on the importance score matrix and the first similarity; and send the source prediction model and the first scaled importance score matrix to the second base station. Thus, the second base station is configured to use the source prediction model and the first scaled importance score matrix to generate a target prediction model and predict radio system parameters related to the second base station. The radio system parameters include future values of user data traffic passing through the second base station. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] The text and drawings are provided as examples only to help the reader understand the present invention. They are not intended to and should not be interpreted as limiting the scope of the present invention in any way. Although certain embodiments and examples have been provided, it will be clear to those skilled in the art based on the disclosure herein that the embodiments and examples shown may be modified without departing from the scope of the embodiments of the present disclosure provided herein.
[0007] Figure 1A The following illustrates a logical flow of determining a target model 1 - 3 based on a similarity 1 - 1 according to some embodiments of the present disclosure.
[0008] Figure 1B An algorithm flow for collecting statistical data, determining source models 1 - 2 , and determining target models 1 - 3 according to some embodiments of the present disclosure is shown.
[0009] Figure 1C The problem 1-80 of poor traffic forecasts and a solution 1-81 to improve forecasts according to some embodiments of the present disclosure are shown.
[0010] Figure 1D Delayed operation parameter selections 1-61 and on-time parameter selections 1-65 according to some embodiments of the present disclosure are compared.
[0011] Figure 2 A system 2-50 according to some embodiments of the present disclosure is shown, which includes a server 2-20 and a base station having a statistical data storage device and a user equipment (UE) generating traffic.
[0012] Figure 3 An example logic flow for determining a target model 1-3, configuring a target base station 3-9, and improving a system 2-50 according to some embodiments of the present disclosure is shown.
[0013] Figure 4 Illustrative hopping diagrams 4-9 illustrating communications and configurations between entities of a system 2-50 are shown, according to some embodiments of the present disclosure.
[0014] Figure 5 Exemplary logic 5 - 9 of a server 2 - 20 is shown, according to some embodiments of the present disclosure.
[0015] Figure 6 Exemplary logic 6-9 is shown that may be performed by a base station of system 2-50 according to some embodiments of the present disclosure.
[0016] Figure 7A and Figure 7B The traffic history and the distribution of the traffic history (probability density or histogram) are shown.
[0017] Figure 8AAn exemplary similarity network 4 - 40 is shown according to some embodiments of the present disclosure.
[0018] Figure 8B and Figure 8C An autoencoder 8-1 and a latent feature vector 8-3 are shown according to some embodiments of the present disclosure.
[0019] Figure 9 It is shown that the solution of the embodiments according to some embodiments of the present disclosure is applied to a newly deployed base station.
[0020] Figure 10 The application of the solution according to some embodiments of the present disclosure in a business demand scenario related to customer service is shown.
[0021] Figure 11 Exemplary hardware for implementing computing devices such as the server 2-20 and the base station of the system 2-50 is shown according to some embodiments of the present disclosure.
[0022] Figure 12 A server according to an embodiment of the present disclosure is shown.
[0023] Figure 13 A base station according to an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0024] Other technical features may be apparent to those skilled in the art from the following drawings, descriptions, and claims.
[0025] Before proceeding to the following detailed description, it may be advantageous to set forth the definitions of certain words and phrases used throughout this patent document. The term "connect" and its derivatives refer to any direct or indirect communication between two or more elements, regardless of whether these elements are in physical contact with each other. The terms "send," "receive," and "communicate" and their derivatives encompass direct and indirect communication. The terms "comprise" and "include" and their derivatives mean, but are not limited to. The term "or" is inclusive, meaning and / or. The phrase "associated with..." and its derivatives mean to include, be included within, interconnect with, include, be included within, be connected to, or be connected to, be coupled to, or be coupled to, can communicate with, collaborate with, interlace, be in parallel, be adjacent to, be coupled to, or be combined with, have, have characteristics, have a relationship with, or have a relationship with. The term "controller" refers to any device, system, or part thereof that controls at least one operation. Such a controller can be implemented in hardware or a combination of hardware and software and / or firmware. Whether local or remote, the functions associated with any particular controller can be centralized or distributed. The phrase "at least one of," when used with a list of items, means that different combinations of one or more of the listed items may be used, and only one item in the list may be required. For example, "at least one of A, B, and C" includes any of the following combinations: A, B, C, A and B, A and C, B and C, and A, B, and C.
[0026] In addition, the various functions described below can be implemented or supported by one or more computer programs, each of which is formed of a computer-readable program code and embodied in a computer-readable medium. The terms "application" and "program" refer to one or more computer programs, software components, instruction sets, processes, functions, objects, classes, instances, related data, or parts thereof that are suitable for implementation with suitable computer-readable program code. The phrase "computer-readable program code" includes any type of computer code, including source code, object code, and executable code. The phrase "computer-readable medium" includes any type of medium that can be accessed by a computer, such as a read-only memory (ROM), random access memory (RAM), a hard drive, a compact disc (CD), a digital video disc (DVD), or any other type of memory. "Non-transitory" computer-readable medium excludes wired, wireless, optical, or other communication links that transmit instantaneous electrical or other signals. Non-transitory computer-readable media include media that can permanently store data, as well as media that can store data and then rewrite data, such as rewritable optical discs or erasable memory devices.
[0027] Definitions for certain other words and phrases are provided throughout this patent document. Those of ordinary skill in the art should understand that in many, if not most, instances, such definitions apply to prior as well as future uses of such defined words and phrases.
[0028] According to an embodiment of the present disclosure, a server is provided that is configured to manage traffic prediction model transfer learning between cellular communication base stations (including 5G base stations as a non-limiting example). The server includes: one or more processors; and one or more memories storing a program, wherein the execution of the program by the one or more processors is configured to cause the server to at least: receive a first plurality of base station statistics, wherein the first plurality of base station statistics includes a first data set of a first size from a first base station; receive a second plurality of base station statistics, wherein the second plurality of base station statistics includes a second data set of a second size corresponding to a second base station; select the first base station as a source base station; train a similarity network; receive a source prediction model and a first importance score matrix from the first base station; receive a prediction model request from a target base station, wherein the target base station is the second base station; calculate a first similarity using the similarity network; obtain a first scaled importance score moment based on the importance score matrix and based on the first similarity; and send the source prediction model and the first scaled importance score matrix to the second base station. Therefore, the second base station is configured to use the source prediction model and the first scaled importance score matrix to generate a target prediction model and predict radio system parameters related to the second base station. The radio system parameters include future values of user data traffic passing through the second base station.
[0029] Problems may arise in a radio communication system when a base station (BS) has a poor configuration relative to the current traffic demand. The base station may not be able to provide the spectrum bandwidth requested by the user equipment device (UE). Or, the base station may be inefficiently allocated more spectrum than is needed to meet the demand. Predictions can be used to determine the spectrum to be allocated to the base station before the demand arises. To reduce prediction errors, statistical models and spatial models in the time domain can use correlations between different base stations (BSs) to make better predictions. However, sharing historical data from one base station to another requires inter-BS bandwidth and memory at the destination base station (BS). In addition, the historical data may be irrelevant to the destination BS. For example, training the destination BS based on historical data may result in learning useless prediction patterns (overtraining on historical data). Important parameters in the source model can be significantly changed by training with a small amount of data from the destination base station.
[0030] At least one embodiment can address these issues by determining a source base station and its similarity to one or more target base stations. Furthermore, the importance of parameters is determined and training is adjusted to account for these importances. Lack of historical data can be compensated for by selecting a base station with extensive historical data as the source base station.
[0031] In some embodiments, execution of the program by one or more processors is further configured to cause the server to: receive a third data set of a third size from a third base station; determine a second similarity using the similarity network and the third data set; calculate a second scaled importance score matrix based on the second similarity and the importance score matrix; and send the source prediction model and the second scaled importance score matrix to the third base station.
[0032] In some embodiments, the first data set includes a histogram of the service history of the source base station, wherein the horizontal axis of the histogram is proportional to the number of bits per second, the first data set further includes a first indication of the frequency bands supported by the source base station, a second indication of the radio access types supported by the source base station, a third indication of the 5G category types supported by the source base station, and a fourth indication of the user density currently supported by the source base station, and wherein a first node vector is formed based on the first data set.
[0033] In some embodiments, execution of the program by the one or more processors is further configured to cause the server to select the candidate base station with the largest data set as the source base station.
[0034] In some embodiments, the similarity network includes an autoencoder, and execution of the program by the one or more processors is further configured to cause the server to train the similarity network by updating parameters of the autoencoder based on an autoencoder loss using gradient descent, wherein the autoencoder loss is a distance between the first node vector and the estimated node vector, wherein the estimated node vector is an output of the similarity network.
[0035] In some embodiments, execution of the program by one or more processors is further configured to cause the server to calculate the first similarity in the following manner: obtaining a second node vector from the target base station; when the first node vector is input to the autoencoder, obtaining a first latent vector as a first output of the autoencoder; when the second node vector is input to the autoencoder, obtaining a second latent vector as a second output of the autoencoder; and calculating the first similarity as a cosine similarity between the first latent vector and the second latent vector.
[0036] In some embodiments, the importance score matrix is a second-order derivative of the Fisher information matrix with respect to the weights of the source prediction model, and the first scaled importance score matrix is a product of the importance score matrix and the first similarity.
[0037] According to another embodiment of the present disclosure, a method for managing business prediction model transfer learning between 5G base stations is provided. The method includes receiving a first plurality of base station statistics, wherein the first plurality of base station statistics include a first data set of a first size from a first base station; receiving a second plurality of base station statistics, wherein the second plurality of base station statistics include a second data set of a second size corresponding to a second base station; selecting the first base station as a source base station; training a similarity network; receiving a source prediction model and a first importance score matrix from the first base station; receiving a prediction model request from a target base station, wherein the target base station is the second base station; calculating a first similarity using the similarity network; obtaining a first scaled importance score matrix based on the importance score matrix and based on the first similarity; and sending the source prediction model and the first scaled importance score matrix to the second base station.
[0038] The present invention also provides a non-transitory computer-readable medium configured to store a program, wherein execution of the program by one or more processors of a server is configured to cause the server to at least: receive a first plurality of base station statistics, wherein the first plurality of base station statistics include a first data set of a first size from a first base station; receive a second plurality of base station statistics, wherein the second plurality of base station statistics include a second data set of a second size corresponding to a second base station; select the first base station as a source base station; train a similarity network; receive a source prediction model and a first importance score matrix from the first base station; receive a prediction model request from a target base station, wherein the target base station is a second base station; calculate a first similarity using the similarity network; obtain a first scaled importance score moment based on the importance score matrix and based on the first similarity; and send the source prediction model and the first scaled importance score matrix to the second base station, whereby the second base station is configured to use the source prediction model and the first scaled importance score matrix to generate a target prediction model and predict radio system parameters related to the second base station, wherein the radio system parameters include future values of user data services passing through the second base station.
[0039] Embodiments of the present disclosure provide TLP, a prediction framework based on transfer learning. Embodiments of the present disclosure achieve high accuracy on traffic prediction tasks with limited and unbalanced data. Embodiments of the present disclosure provide data and bandwidth efficiency. Embodiments of the present disclosure use available data from data-rich reference nodes (base stations) and data-limited regular nodes (base stations) without migrating traffic history records.
[0040] Embodiments of the present disclosure use a relatively large amount of data from reference edge nodes (also called source base stations) to train a basic neural network (NN) model (called a source model or source prediction model), which not only extracts node-specific features but also (to some extent) summarizes general features.
[0041] Embodiments of the present disclosure use a limited amount of data maintained by a target base station to fine-tune the weights of a source model to serve another node (target base station).
[0042] A major challenge in migrating prediction models between edge nodes is the difficulty in maintaining general features while updating node-specific features with a small amount of data at the edge (target base station). Embodiments of the present disclosure use layer freezing (LF), elastic weight curing (EWC), or similarity-based EWC (SEWC) to maintain general features while updating node-specific features.
[0043] The embodiments of the present disclosure are more fine-grained and node-customized than previous approaches. The embodiments of the present disclosure provide weight-level importance scores to balance between 1) preserving learned NN weights for generality and 2) updating those weights to serve specific goals. The embodiments of the present disclosure also customize model migration for each individual edge node by adjusting the importance score based on the similarity between the source and target.
[0044] Figure 1A Exemplary logic 1-12 for determining a target model 1-3 according to an embodiment of the present disclosure is shown. At operation 1-10, a similarity 1-1 is determined. At operation 1-11, the source model 1-2 is modified using the similarity 1-1 to provide a target model 1-3.
[0045] Figure 1B An exemplary algorithm flow 1-3 according to an embodiment of the present disclosure is shown. In algorithm state 1, statistical data are collected into a data set at a base station. These statistical data, such as data set 1-21 from base station (BS) 1-31 and data set 1-22 from BS 1-32, are input to algorithm state 2. In algorithm state 2, a source model 1-2 is determined, and a similarity 1-1 is determined (e.g., between BS 1-31 and BS 1-32). In algorithm state 3, transfer learning is performed by training a target model 1-3.
[0046] Figure 1C Problems 1-80 and solutions 1-81 of business forecasting according to embodiments of the present disclosure are shown.
[0047] exist Figure 1C At the top of , curve 1-51 shows how traffic starts increasing in the early morning (4 a.m.), reaches a peak (around 6 p.m.) and tapers off (midnight).
[0048] exist Figure 1CAt the bottom of the diagram, the allocated spectrum is shown with a vertical axis 1-52 and a horizontal time axis 1-53. Past allocations are represented by 1-70, and future allocations are represented by 1-71. The vertical line 1-50 indicates the time at which the forecast is performed. If the future allocation indicates a downward trend in traffic, a problem 1-80 may arise. As a non-limiting example, embodiments of the present disclosure provide a solution 1-81 in which the forecast more closely follows the future trend. Figure 1C is a schematic diagram for explaining a communication system 2-50 ( Figure 2 )'s example problems and solutions. The embodiments of the present disclosure are not limited to radio systems.
[0049] Figure 1D It further shows Figure 1C Questions 1-80 and Figure 1C Solution 1-81. Time progresses from top to bottom, as shown by timeline 1-69. On the left, events 1-60, 1-61, and 1-62 shown at 10 o'clock, 11 o'clock, and 12 o'clock are shown without prediction and result in delayed operation (for example, the allocated spectrum is insufficient to meet demand, causing some user equipment (UE) to wait for service). On the right are events 1-63, 1-64, 1-65, and 1-66 with predictions at 10 o'clock, 11 o'clock, and 12 o'clock. The output of operation 1-64 is the predicted traffic at the future time 12 o'clock. Resources have been allocated appropriately, and at 1-66 the resources are sufficient to meet demand, and the number of UEs waiting for service is reduced.
[0050] Figure 2 An exemplary system 2-50 according to an embodiment of the present disclosure is shown. The number of entities is symbolic and is not to scale with the number in a working system. BS 2-1 stores statistical data 2-21 in data set 2-31. A dashed line is shown as associating BS 2-1 with the storage device. Typically, the storage device of the BS is geographically local to the BS (rather than geographically remote from the BS). BS 2-2 also stores statistical data 2-22 in data set 2-32. Similarly, statistical data 2-23 in data set 2-33 at BS 2-3 and statistical data 2-24 in data set 2-34 at BS 2-4. Figure 2 Also shown are UEs 2-10, 2-11, 2-12, 2-13, and 2-14. Figure 2 The base station and UE are Figure 1B An example of a base station and a UE.
[0051] BS 2-1 experiences traffic 2-15, and BS 2-2 experiences traffic 2-16. Symbol Ω may be used herein to represent the history of traffic.
[0052] Figure 3Exemplary logic 3-8 for determining a target model 1-3 and improving the operation of a system 2-50 is shown according to an embodiment of the present disclosure.
[0053] At operation 3-30, a source base station 3-1 is selected. At operation 3-32, a similarity network 3-3 is trained based on first base station statistics 3-5 of the source base station 3-1 and second base station statistics 3-7 of the target base station 3-9. At operation 3-34, the similarity network 3-3 is used to determine a similarity 1-1 between the source base station 3-1 and the target base station 3-9. At operation 3-36, a scaled importance score matrix 3-13 is determined based on the importance score matrix 3-11 of the source base station 3-1 and the similarity 1-1.
[0054] Then, at operation 3-38, the scaled importance score matrix 3-13 and the source model 1-2 are sent to the target base station 3-9. At operation 3-40, the target base station 3-9 determines the target model 1-3. The target base station 3-9 then uses the target model 1-3 to form a prediction 3-16 of the local traffic 2-16 at a future time 3-19. The target base station 3-9 is then configured with an allocation 1-71 at the appropriate time to support the traffic 2-16 at time 3-19.
[0055] Figure 4 The embodiment of the present disclosure corresponds to Figure 3 4-9 shows an exemplary jump diagram of the logic of the system. A timeline progressing from top to bottom is shown on the left. At the top, the system entities of server 2-20, source base station 3-1, target base station 3-9, UE 2-11, and UE 2-13 are shown. The ellipsis shown indicates that there are typically many base stations and many UEs. There may be several source base stations, each providing source models to many target base stations. A base station is selected as a source in part because the traffic at the source is similar to the traffic at another base station for which a model is needed.
[0056] exist Figure 4 , source base station 3-1 provides data set 4-1 to server 2-20 (message 4-10), and target base station 3-9 provides data set 4-3 and model request 4-4 to server 2-20 (message 4-12, which may be several messages). After receiving data set 4-1 and data set 4-3, server 2-20 may determine which base station is the source base station and which base station is the target base station.
[0057] At operation 4-14, the server 2-20 trains the similarity network 3-3, including training the similarity network 4-40. At operation 4-15, the server 2-20 determines the similarity 1-1 between the source base station 3-1 and the target base station 3-9.
[0058] At operation 4-16, the source base station 3-1 determines the source model 1-2 and the importance score matrix 3-11. This may be caused by a command from the server 2-20 (the command is not shown). In message 4-17 (which may be several messages), the source base station 3-1 sends the importance score matrix 3-11 and the source model 3-15 to the server 2-20.
[0059] Then, at operation 4-18, the server 2-20 determines the scaled importance score matrix 3-13 and sends it along with the source model 1-2 to the target base station 3-9 in a message 4-20.
[0060] At operation 4-22, the target base station 3-9 generates a target model 1-3 using the scaled importance score matrix 3-13, the source model 1-2, and the statistical data 2-22 (statistical data local to the target base station 3-9).
[0061] At operations 4-24 and 4-25, the target base station 3-9 and the source base station 3-1 form respective predictions 3-16 and 4-34 for the respective services 2-16 and 2-15. Then, at operation 4-26, the target base station 3-9 is configured based on the prediction 3-16. At operation 4-27, the source base station 3-1 is configured based on the prediction 4-34.
[0062] At time 3-19, actual traffic 4-30 occurs (labeled 4-28). Performance 3-20 is improved, and costs 3-21 are reduced (labeled 4-32).
[0063] Figure 5 Exemplary server logic 5 - 9 is shown that is executed by a server 2 - 20 according to an embodiment of the present disclosure.
[0064] The server 2-20 determines which base stations are sources, which base stations are targets, and the source base station for each target base station.
[0065] At operation 5-30, the server 2-20 collects statistical data of the BS 2-1, BS2-2, and other base stations.
[0066] At operation 5-32, the server selects a base station as the source base station 3-1 based in part on the base station having a large number of data points stored. This number of data points may correspond to, for example, several (3 to 12) months of traffic history records. A small number of data points would correspond to, for example, several weeks (2 to 6 weeks) of traffic history records.
[0067] After operation 5-32, two logic flows occur, one logic flow starts at operation 5-34, and the other logic flow starts at operation 5-46. These can be done in parallel or in series.
[0068] At operation 5-34, the server 2-20 trains a similarity network 4-40 based on the statistical data 3-5 of the source base station 3-1. At operation 5-36, the server 2-20 receives a model request 4-4 from the target base station 3-9. At operation 5-38, the server 2-20 calculates the source latent features 5-3 and the target latent features 5-5 using the similarity network 4-4.
[0069] At operation 5-40, the server 2-20 calculates the similarity 1-1 between the target base station 3-9 and the source base station 3-1.
[0070] Turning to operation 5-46, the server 2-20 sends a request 5-5 to the source base station 3-1, asking it to train the source model 1-2. At operation 5-48, the server 2-20 receives the source model 1-2 and the importance score matrix 3-11 from the source base station 3-1.
[0071] The parallel logic flows converge at operation 5-42, and the server 2-20 calculates the scaled importance matrix 3-13 as the product (e.g., Kronecker product) of the similarity 1-1 (scalar) and the importance score matrix 3-11. Then, at operation 5-44, the server 2-20 sends the source model 1-2 and the scaled importance score matrix 3-13 to the target base station 3-9.
[0072] Figure 6 1 shows exemplary base station logic 6-9 according to an embodiment of the present disclosure. At the beginning of the logic flow, the base station does not know whether it is a source base station or a target base station. This determination is made by the server 2-20.
[0073] At operation 6-30, the base station collects a history 6-1 of traffic 6-3 and determines a histogram 6-5. At operation 6-32, the base station forms a data set 6-7 including base station statistics 6-8. Then, at operation 6-34, the base station uploads the data set 6-7 to the server 2-20.
[0074] At operation 6-36, the base station determines whether a model request 5-1 has been received from the server 2-20. This determination can be made periodically or on some other predetermined schedule. If so, the base station is the source base station and the logic of operation 6-38 is then executed, i.e., the source model 1-2 is trained. If not, the base station is the target base station and the logic of operation 6-50 is then executed.
[0075] Please refer to operation 6-50. The base station is a target base station, for example, the target base station 3-9. The target base station 3-9 sends a model request 4-4 to the server 2-20. At operation 6-52, the target base station 3-9 downloads the source model 1-2 and the scaled importance score matrix 3-13 from the server 2-20. At operation 6-54, the target base station 3-9 determines the target model 1-3. At operation 6-56, the target base station 3-9 forms a prediction 3-16 of the service 2-16. At 6-58, the target base station 3-9 is configured based on the prediction 3-16. The spectrum allocation of the BS is configured at the time of deployment. The spectrum allocation is fixed at that time. If the spectrum allocation is subsequently changed by some dynamic algorithm, the server according to an embodiment of the present disclosure triggers the similarity calculation and similar operations as described herein. At operation 6-60, the target base station 3-9 supports the service 2-16.
[0076] Please refer to operation 6-36. If the source model request 5-1 is received at a base station, the logic flows to operation 6-38, and the base station is a source base station, such as source base station 3-1. Source base station 3-1 performs two actions, and these actions can be performed in parallel. At operation 6-40, source base station 3-1 calculates importance score matrix 3-11. At operation 6-42, source base station 3-1 uploads source model 1-2 and importance score matrix 3-11 to server 2-20. Server 2-20 can then provide source model 1-2 to another base station, see operation 6-52.
[0077] Referring again to operation 6-38, the source base station 3-1 also supports services local to the source base station 3-1. At operation 6-46, the source base station 3-1 uses the source model 1-2 to form a prediction 4-34 for service 2-15. At operation 6-48, the source base station 3-1 is configured based on the prediction 4-34. At 6-49, the source base station 3-1 supports service 2-15.
[0078] Figure 7A A graph 7-1 including a traffic history Ω6-3 in a time series according to an embodiment of the present disclosure is shown. An x-axis 7-4 is time, and a y-axis 7-3 shows the intensity of traffic normalized by a maximum value.
[0079] Figure 7B Graph 7-11 is shown including a normalized density or histogram of Ω 6-3, 6-5, according to an embodiment of the present disclosure. The x-axis 7-14 is the random variable of the normalized traffic, and the y-axis 7-13 is the probability of a given random variable value.
[0080] Figure 8AA system 8-9 including a similarity network 4-40 and a similarity calculator 8-11 according to an embodiment of the present disclosure is shown. The similarity network 4-40 includes an encoder 8-1 and a decoder 8-5. The first base station statistics 3-5 (source) G_S are input to the encoder 8-1, and a latent vector Z_S5-5 is obtained. The second base station statistics 3-7 (target) G_T are input to the encoder 8-1, and a latent vector Z_T 5-3 is obtained. The similarity calculator 8-11 calculates, for example, cosine similarity and provides a similarity 1-1 (here denoted as μ). Based on the latent vector Z_S5-5, the decoder has been trained to reconstruct the base station statistics, providing estimated first base station statistics 8-7 in the example shown, It is part of the training process of the similarity network 4-40 and is not directly used to determine the similarity 1-1.
[0081] Figure 8B The general operation 8-50 of the encoder 8-1 is shown, accepting input 6-9, "G", and producing a latent feature vector 8-3, "Z". Figure 8C A schematic representation of the latent feature vector Z shown in FIG8-51 is shown, with the x-axis being 8-57 and the y-axis being 8-58.
[0082] Figure 9 1 shows a transition from a reference cellular communication BS 9-1 to a newly deployed cellular communication BS 9-2 (as a non-limiting example, these may be Figure 9 The model migration 9-9 of the 5G base station mentioned in the embodiment of the present invention is performed. The source model 1-2 is migrated using logic 3-8 to obtain the target model 1-3. Then, operation 3-42 is performed to predict the business, and then operation 3-44 is used to configure the target base station 9-2.
[0083] Figure 10 The application of logic 3-8 to customer service demand forecasting according to an embodiment of the present disclosure is shown in Figure 10-9. A source model has been established in a first city 10-1 (e.g., Montreal, Quebec, Canada), where a person 10-2 already provides customer service (e.g., via telephone or internet communication such as email or chat software). A customer service operation is established in a second city 10-6 (e.g., New York, New York, USA), where a person 10-4 will provide service. Logic 10-3 (modifying 3-8 as needed) is used to provide allocation of resources 10-5 based on the forecast.
[0084] Now available for execution Figures 3 to 6 An example method of the logic.
[0085] Each base station records traffic data samples. The data set from the source base station is denoted as S_S.
[0086] The traffic sample s[t] can be written as equation (1). a[t] is the traffic volume.
[0087] s[t]=[a[t];t] Equation (1)
[0088] A window c of traffic samples can be formed for time series analysis as in equation (2).
[0089] x[t]=[a[t],a[t+1],…,a[t+c-1]; t,t+1,…,t+c-1] Equation (2)
[0090] as well as
[0091] y[t]=a[t+c] Equation (3)
[0092] The transformed data sample is u[t] = {x[t]; y[t]}. The window is moved forward one sample to generate u[t+1]. The transformed data set is Ω_S = {X_S, Y_S}, where X_S = {x[t]} and Y_S = {y[t]} represent all input vectors and output vectors transformed from S_S. The data set of the k-th target base station is denoted as Ω_T(k). The source models 1-2 are denoted as θ_B below. The target models 1-3 of the k-th base station are denoted as θ_T(k) below.
[0093] The prediction loss of the new data sample x[t] at the target base station 3-9 can be defined as in equation (4).
[0094]
[0095] where the sum (“Σ”) is over y[t]∈Ω T |Ω| is the size of the dataset, is the prediction of the ground truth y, and d(.,.) is the error metric (e.g., absolute error, squared error, root mean square error).
[0096] Embodiments of the present disclosure provide target models 1-3 (θ_B) for small Ω_T from a large dataset Ω_S while minimizing the loss L of equation (4) for target base stations 3-9.
[0097] One model transfer technique is weight initialization. The weights start with values w i,j [0] and then updated using stochastic gradient descent (SGD).
[0098]
[0099] Here, t represents the period number.
[0100] When training is complete, the model θ_T has been learned.
[0101] Another migration technique is layer freezing, as shown in the next two rows, which form equation (6).
[0102] w i,j [t+1]=w i,j [t]ifj≤β
[0103] For other j, update using
[0104] Another migration technique is to use elastic weight curing of the Fisher Information Matrix (FIM). Let Fi,j be the diagonal value of the FIM of wi,j. Fi,j can be calculated using the first-order derivative as shown in Equation (7).
[0105]
[0106] where the expected values ("E") exceed those (x,y) in Ω_S.
[0107] Fi,j is now a measure of w i,j A constant score for importance.
[0108] Given this importance score, the loss function for transferring and / or fine-tuning the neural network (NN) at the target is now updated as in Equation (8).
[0109] L EWC =L+λL R =L+λΣ0.5F i,j (w i,j -w i,j [0]) 2 Equation (8)
[0110] where the sum exceeds i,j, wi,j at t=0 is the initial value of wi,j before migration, and the initial weights in the copied model (here called θ_D) and λ control the balance between the prediction loss and the regularization term associated with Fi,j.
[0111] The weight of the transferred model θ_T (target model 1-3) is updated as in equation (9).
[0112]
[0113] Among them, at a given period t,
[0114] Embodiments of the present disclosure customize the importance score for each individual target. Ideally, if the target is very similar to the source, we tend to keep the importance score high to maintain the common characteristics obtained from the source. Otherwise, we lower the importance score so that the transferred model is not overly constrained by the base model.
[0115] The embodiment of the present disclosure provides a similarity-based EWC (SEWC) technology. SEWC uses an automatic encoder (AE) (e.g., Figure 8A The similarity network 4-40) describes the similarity between the source node and the target node in the high-level latent space. Figure 8A As shown (and also discussed above), the input to the similarity AE is a vector G that describes the characteristics of the node (target or source).
[0116] In some embodiments, to describe a base station (also referred to herein as a node), the following information (also referred to herein as indications) is concatenated: 1) traffic distribution, 2) supported 3GPP frequency bands (including 3G, LTE / 4G, and 5G frequency bands), 3) supported radio access technology (RAT) types (including GSM, UMTS, FDD-LTE, TDD-LTE, and 5G NR), 4) the category of 5G New Radio (NR) (if applicable) (i.e., wide area, medium range, or local area node), and / or 5) the user density level around the node. The collection of indications may be referred to herein as a data set and statistics. These statistics / indications will be further described below.
[0117] Traffic distribution H: It is represented by a 100-bin normalized histogram of the node's historical traffic, as shown in equation (10).
[0118] H = histo(Ω) Equation (10)
[0119] Among them, histo is the calculation of the normalized histogram.
[0120] The supported 3GPP frequency bands FB are represented by a binary vector,
[0121] F B =f b (1),f b (2),…,f b (N b )wheref b (k)=1 indicates that the described base station supports the kth frequency band, and NB is the total number of frequency bands defined in the 3GPP standard. The frequency bands are connected in the order of 3G band, LTE / 4G band and 5G band.
[0122] The supported radio access technology (RAT) type F_RAT is represented by a binary 5-dimensional vector,
[0123] F RAT =[f RAT (1),…,f RAj (5)],wheref RAT (k) = 1 indicates that the described base station supports the kth RAT type. RAT types are arranged in the order of GSM, UMTS, FDD-LTE, TDD-LTE, and 5G NR. For example, F_RAT = [0; 0; 1; 0; 1] means that the current node supports FDD-LTE and 5G NR.
[0124] A type of 5G NR F_5G uses a binary 5-dimensional vector F 5G =[f 5G (1),…,f 5G (4)], where f_5G(k) = 1 indicates that the current node is a class k 5G node. The classes are organized in the order of wide area, medium range, and local area. For example, F_5G = [1; 0; 0] means that the current node is a wide area 5G NR node. For further description of the classes, see TS 138 104-V15.3.0 - 5G; NR; Base Station (BS) Radio Transmission and Reception (3GPP TS 38.104 Version 15.3.0 Release 15).
[0125] Describes the user density level ρ around the BS. ρ is a binary 3-dimensional vector ρ = [ρ(1), ρ(2), ρ(3)]. Density levels are arranged in the order of high, medium, and low. For example, ρ = [1; 0; 0] indicates that the user density around the current node is high. Some examples of user density thresholds are as follows: Th_low: 1,000 people per square kilometer, Th_high: 10,000 people per square kilometer. If the density is greater than Th_high, the density is high; otherwise, if the density is less than Th_low, the density is low; otherwise, the density is medium.
[0126] The above indicators or vectors are connected as in equation (11).
[0127] G=[H,F B ,F RAT ,F 5G ,ρ] Equation (11)
[0128] like Figure 8A As shown, the encoder 8-1 of the similarity AE (also referred to herein as the similarity network 4-40) first converts the vector G ( Figure 8B Items 6-7) are mapped to the vector of latent features z ( Figure 8A z then passes through the decoder 8-5 of the AE to generate the replica Its goal is to replicate G 6-7 with high similarity (an exact copy will have the highest similarity).During the training process, the AE uses only the information of the source nodes G_S 3-5 and learns to minimize the reconstruction loss LRECON as defined in Equation (12).
[0129]
[0130] After training the AE, the encoder 8-1 is used to generate a latent feature (called Z 8-3) from the source node description GS of the source base station 3-1 and the target node description GT of the target base station 3-9. The corresponding latent features are z_S 5-5 and z_T 5-3 (see Figure 8A ). The similarity calculator calculates the cosine similarity between z_S and z_T, and the result 1-1 (also called μ) is used as the source-target similarity score. The calculation is shown in equation (13).
[0131]
[0132] This similarity score (μ, also called similarity 1-1) is then multiplied by the importance score of EWC to generate a node similarity-aware importance score as shown in Equation (14).
[0133] FSIM i,j =μF i,j Equation (14)
[0134] The loss function for transferring the base model is given in Equation (15).
[0135] L SEWC =L+λΣ0.5μF i,j (w i,j -w i,j [0]) 2 Equation (15)
[0136] Among them, the sum exceeds i,j.
[0137] The weights are updated at each epoch as shown in Equation (16).
[0138]
[0139] The above techniques range from the entire model level (for WI), the separation layer level (for LF), to the individual weight level (for EWC and SEWC). WI treats all weights in the model indifferently; LF distinguishes between the layers based on the generality of the features extracted; and EWC and SEWC assign different importance scores to individual weights.
[0140] In order to integrate them all into a single transfer learning-based prediction (TLP) framework, the embodiments of the present disclosure provide a unified formula. The embodiments of the present disclosure achieve this by reformulating the weight update rules of these techniques.
[0141] Let η_(i, j) be the learning rate of weight wi,j during the model transfer process. In each training epoch, weight wi,j follows the uniform update rule given by Equation (17).
[0142]
[0143] If the WI technique is applied to model transfer, the learning rate remains the same for all weights, as shown in Equation (18).
[0144]
[0145] Among them, η * Can be a constant or a variable that changes based on the training epochs (but remains constant for all weights within an epoch).
[0146] If the LF technique is selected, different learning rates are used for the front and back layers, as shown in Equation (19) in the following two rows.
[0147]
[0148] Regarding the learning rate of EWC, the first step is to take the partial derivative of Equation (8) with respect to the weights wi,j. Combining the result with Equation (17), the learning rate for all i,j is given by Equation (20).
[0149]
[0150] For SEWC, the learning rate is similar to Equation (20) and is given in Equation (21).
[0151]
[0152] If the base model is directly applied to the target without any training, the learning rate is zero for all i,j, as shown in Equation (22).
[0153] η i,j =0 Equation (22)
[0154] Considering Equation (17), the essence of transfer learning can be explained.
[0155] As provided in the embodiments of the present disclosure, SEWC balances the need to preserve general features and learn new features by adding a similarity-based regularization term to the update process of each weight in the model. Increasing the regularization term helps better preserve general features. Reducing the regularization term allows the model to adapt more freely to new target data.
[0156] Substituting equation (21) into equation (17) provides equation (23).
[0157] w i,j [t+1]=w i,j [t]-η * λμF i,j (w i,j -w i,j [0]) Equation (23)
[0158] In equation (23), the second term on the right is the regular weight increment caused by the prediction error gradient, and the third term is the additional weight increment brought by SEWC. Let this SEWC weight increment be δ i,j , as shown in equation (24).
[0159] δ i,j =-η * λμF i,j (w i,j -w i,j [0]) Equation (24)
[0160] If w i,j =w i,j [0], then δ i,j = 0. In this case, the weights remain unchanged from their original values. Therefore, SEWC does nothing.
[0161] If w i,j >w i,j [0], then δ i,j <0. In this case, the weight is greater than its original value. SEWC applies negative δ i,j to reduce the weight. A higher similarity score μ or a larger importance score Fi,j will bring a larger decrease and thus assign wi,j a stronger i,j Retweet of [0].
[0162] If w i,j <w i,j [0], then δ i,j > 0. In this case, the weight is smaller than its original value. If the similarity score μ is high or the weight is important (large Fi,j), SEWC imposes a positive δ i,j to make the weights reach their original values faster.
[0163] From another perspective, Fi,j also controls the sensitivity of wi,j to the target data. When wi,j is updated by new data, if μ or Fi,j is large, it will cause high loss.
[0164] Example quantified benefits are as follows.
[0165] use Figure 3 The similarity 1-1 in the logic of (see also Eq. (23)) allows for better performance than the next best approach (i.e., not using Figure 8A A neural network (NN) with a similarity network of 4-40 predicts load more quickly and very accurately. NN is a predictive neural network that uses pattern recognition and machine learning. Furthermore, embodiments of the present disclosure provide better load forecasts at a given time than NN (with 20% greater accuracy) and also provide better load forecasts than the autoregressive integrated moving average (ARIMA) model used for forecasting time series (with approximately 50% greater accuracy).
[0166] Now refer to Figure 11 Hardware is described for performing the embodiments of the disclosure provided herein.
[0167] Figure 11 An exemplary device 11-1 for implementing the embodiments disclosed herein is shown. For example, the device 11-1 can be a server, a computer, a laptop computer, a handheld device, or a tablet computer device. The device 11-1 can include one or more hardware processors 11-9. The one or more hardware processors 11-9 can include an ASIC (application-specific integrated circuit), a CPU (e.g., a CISC or RISC device), and / or custom hardware. The device 11-1 can also include a user interface 11-5 (e.g., a display screen and / or a keyboard and / or a pointing device such as a mouse). The device 11-1 can include one or more volatile memories 11-2 and one or more non-volatile memories 11-3. The one or more non-volatile memories 11-3 can include a non-transitory computer-readable medium that stores instructions that are executed by the one or more hardware processors 11-9 to cause the device 11-1 to perform any method of the embodiments disclosed herein.
[0168] Figure 12 A server according to an embodiment of the present disclosure is schematically illustrated.
[0169] refer to Figure 12 , the server 1200 may include a processor 1210, a transceiver 1220, and a memory 1230. However, not all of the components shown are required. The server 1200 may be composed of Figure 12The components shown may be implemented with more or fewer components. In addition, according to another embodiment, the processor 1210, the transceiver 1220, and the memory 1230 may be implemented as a single chip.
[0170] The above-mentioned components will now be described in detail.
[0171] The processor 1210 may include one or more processors or other processing devices that control the proposed functions, processes, and / or methods. The operations of the server 1200 may be implemented by the processor 1210.
[0172] The transceiver 1220 may include an RF transmitter for up-converting and amplifying a transmit signal, and an RF receiver for down-converting the frequency of a receive signal. However, according to another embodiment, the transceiver 1220 may be implemented by more or fewer components than those shown.
[0173] The transceiver 1220 may be connected to the processor 1210 and transmit and / or receive signals. The signals may include control information and data. In addition, the transceiver 1220 may receive signals through a wireless channel and output the signals to the processor 1210. The transceiver 1220 may transmit signals output from the processor 1210 through a wireless channel.
[0174] The memory 1230 may store control information or data included in the signal obtained by the server 1200. The memory 1230 may be connected to the processor 1210 and store at least one instruction or protocol or parameter for the proposed function, process and / or method. The memory 1230 may include a read-only memory (ROM) and / or a random access memory (RAM) and / or a hard disk and / or a CD-ROM and / or a DVD and / or other storage devices.
[0175] Figure 13 A base station according to an embodiment of the present disclosure is schematically illustrated.
[0176] refer to Figure 13 , the base station 1300 may include a processor 1310, a transceiver 1320, and a memory 1330. However, not all of the components shown in the figure are required. The base station 1300 may be composed of Figure 13 Furthermore, according to another embodiment, the processor 1310 , the transceiver 1320 , and the memory 1330 may be implemented as a single chip.
[0177] In an exemplary embodiment, the base station 1300 may be a gNodeB. In an exemplary embodiment, the above-mentioned source BS and target BS may correspond to the base station 1300.
[0178] The above-mentioned components will now be described in detail.
[0179] The processor 1310 may include one or more processors or other processing devices that control the proposed functions, processes, and / or methods. The operations of the base station 1300 may be implemented by the processor 1310.
[0180] The transceiver 1320 may include an RF transmitter for up-converting and amplifying a transmit signal, and an RF receiver for down-converting the frequency of a receive signal. However, according to another embodiment, the transceiver 1320 may be implemented by more or fewer components than those shown.
[0181] The transceiver 1320 may be connected to the processor 1310 and transmit and / or receive signals. The signals may include control information and data. In addition, the transceiver 1320 may receive signals through a wireless channel and output the signals to the processor 1310. The transceiver 1320 may transmit signals output from the processor 1310 through a wireless channel.
[0182] The memory 1330 may store control information or data included in the signal obtained by the base station 1300. The memory 1330 may be connected to the processor 1310 and store at least one instruction or protocol or parameter for the proposed function, process and / or method. The memory 1330 may include a read-only memory (ROM) and / or a random access memory (RAM) and / or a hard disk and / or a CD-ROM and / or a DVD and / or other storage devices.
[0183] Although the present disclosure has been described using exemplary embodiments, various changes and modifications may be suggested to those skilled in the art. The present disclosure is intended to encompass such changes and modifications as fall within the scope of the appended claims. Nothing in this application should be construed as implying that any particular element, step, or function is essential to be included within the scope of the claims. The scope of a patented subject matter is defined by the claims.
Claims
1. A server configured to manage traffic prediction model transfer learning between 5G base stations, the server comprising: one or more processors; as well as One or more memories storing a program, wherein execution of the program by the one or more processors is configured to cause the server to at least: receiving a first plurality of base station statistics, wherein the first plurality of base station statistics comprises a first data set of a first size from a first base station; receiving a second plurality of base station statistics, wherein the second plurality of base station statistics comprises a second data set of a second size corresponding to a second base station; selecting the first base station as a source base station; Train the similarity network; receiving a source prediction model and a first importance score matrix from the first base station; receiving a prediction model request from a target base station, wherein the target base station is the second base station; Calculating a first similarity using the similarity network; obtaining a first scaled importance score matrix based on the importance score matrix and based on the first similarity; as well as sending the source prediction model and the first scaled importance score matrix to the second base station, Thereby, the second base station is configured to generate a target prediction model using the source prediction model and the first scaled importance score matrix and predict radio system parameters related to the second base station, wherein the radio system parameters include future values of user data traffic passing through the second base station.
2. The server of claim 1 , wherein execution of the program by the one or more processors is further configured to cause the server to: receiving a third data set of a third size from a third base station; determining a second similarity using the similarity network and the third data set; calculating a second scaled importance score matrix based on the second similarity and the importance score matrix; as well as The source prediction model and the second scaled importance score matrix are sent to the third base station.
3. A server according to claim 1, wherein the first data set includes a histogram of the service history of the source base station, wherein the horizontal axis of the histogram is proportional to the number of bits per second, the first data set also includes a first indication of the frequency band supported by the source base station, a second indication of the radio access type supported by the source base station, a third indication of the 5G category type supported by the source base station, and a fourth indication of the user density currently supported by the source base station, and wherein a first node vector is formed based on the first data set. 4 . The server of claim 1 , wherein execution of the program by the one or more processors is further configured to cause the server to select a candidate base station with a largest data set as the source base station.
5. The server of claim 3 , wherein the similarity network comprises an autoencoder, and execution of the program by the one or more processors is further configured to cause the server to train the similarity network by updating parameters of the autoencoder based on an autoencoder loss using gradient descent, wherein the autoencoder loss is a distance between the first node vector and an estimated node vector, wherein the estimated node vector is an output of the similarity network.
6. The server according to claim 5, wherein the execution of the program by the one or more processors is further configured to cause the server to calculate the first similarity in the following manner: Obtain a second node vector from the target base station; When the first node vector is input to the autoencoder, a first latent vector is obtained as a first output of the autoencoder; When the second node vector is input to the autoencoder, a second latent vector is obtained as a second output of the autoencoder; as well as The first similarity is calculated as a cosine similarity between the first latent vector and the second latent vector.
7. The server of claim 3, wherein the importance score matrix is a second-order derivative of a Fisher information matrix with respect to weights of the source prediction model, and the first scaled importance score matrix is a product of the importance score matrix and the first similarity.
8. A method for managing traffic prediction model transfer learning between 5G base stations, the method comprising: receiving a first plurality of base station statistics, wherein the first plurality of base station statistics comprises a first data set of a first size from a first base station; receiving a second plurality of base station statistics, wherein the second plurality of base station statistics comprises a second data set of a second size corresponding to a second base station; selecting the first base station as a source base station; Train the similarity network; receiving a source prediction model and a first importance score matrix from the first base station; receiving a prediction model request from a target base station, wherein the target base station is the second base station; Calculating a first similarity using the similarity network; obtaining a first scaled importance score matrix based on the importance score matrix and based on the first similarity; as well as sending the source prediction model and the first scaled importance score matrix to the second base station, Thereby, the second base station is configured to generate a target prediction model using the source prediction model and the first scaled importance score matrix and predict radio system parameters related to the second base station, wherein the radio system parameters include future values of user data traffic passing through the second base station.
9. The method according to claim 8, further comprising: receiving a third data set of a third size from a third base station; determining a second similarity using the similarity network and the third data set; calculating a second scaled importance score matrix based on the second similarity and the importance score matrix; as well as The source prediction model and the second scaled importance score matrix are sent to the third base station.
10. A method according to claim 8, wherein the first data set includes a histogram of the service history of the source base station, wherein the horizontal axis of the histogram is proportional to the number of bits per second, the first data set also includes a first indication of the frequency bands supported by the source base station, a second indication of the radio access type supported by the source base station, a third indication of the 5G category type supported by the source base station, and a fourth indication of the user density currently supported by the source base station, and wherein a first node vector is formed based on the first data set.
11. The method of claim 8, further comprising selecting a candidate base station with a largest data set as the source base station.
12. The method of claim 10, wherein the similarity network comprises an autoencoder, and wherein training the similarity network comprises updating parameters of the autoencoder based on an autoencoder loss using gradient descent, wherein the autoencoder loss is a distance between the first node vector and an estimated node vector, wherein the estimated node vector is an output of the similarity network.
13. The method according to claim 12, wherein the calculating the first similarity comprises: Obtain a second node vector from the target base station; When the first node vector is input to the autoencoder, a first latent vector is obtained as a first output of the autoencoder; When the second node vector is input to the autoencoder, a second latent vector is obtained as a second output of the autoencoder; as well as The first similarity is calculated as a cosine similarity between the first latent vector and the second latent vector.
14. The method of claim 8, wherein the importance score matrix is a second-order derivative of a Fisher information matrix with respect to the weights of the source prediction model, and the first scaled importance score matrix is a product of the importance score matrix and the first similarity.
15. A non-transitory computer-readable medium configured to store a program, wherein execution of the program by one or more processors of a server is configured to cause the server to at least: receiving a first plurality of base station statistics, wherein the first plurality of base station statistics comprises a first data set of a first size from a first base station; receiving a second plurality of base station statistics, wherein the second plurality of base station statistics comprises a second data set of a second size corresponding to a second base station; selecting the first base station as a source base station; Train the similarity network; receiving a source prediction model and a first importance score matrix from the first base station; receiving a prediction model request from a target base station, wherein the target base station is the second base station; Calculating a first similarity using the similarity network; obtaining a first scaled importance score matrix based on the importance score matrix and based on the first similarity; as well as sending the source prediction model and the first scaled importance score matrix to the second base station, Thereby, the second base station is configured to generate a target prediction model using the source prediction model and the first scaled importance score matrix and predict radio system parameters related to the second base station, wherein the radio system parameters include future values of user data traffic passing through the second base station.
Citation Information
Patent Citations
Physical cell identification distribution method based on fuzzy hierarchical clustering
CN109635858A
Procedure for optimization of self-organizing network
WO2020048594A1