COMMUNICATION SYSTEM AND METHOD FOR TRAINING A MACHINE LEARNING MODEL FOR A COMMUNICATION SYSTEM - Patent application
The hierarchical transfer learning approach in communication systems addresses the challenge of training diverse machine learning models by customizing common models at lower levels, improving training efficiency and accuracy.
Patent Information
- Application Number
- JP2024576989
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-27
- Filing Date
- 2024-10-07
- Publication Date
- 2025-12-26
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing communication systems face challenges in efficiently training machine learning models for diverse objectives, coverage areas, and device types due to high heterogeneity, leading to complex models with high training data requirements and low accuracy across different categories.
A hierarchical transfer learning approach is employed, where a common machine learning model is trained at higher levels and customized at lower levels, utilizing multi-level aggregation and filtering to reduce training data overhead and improve accuracy.
This method enhances training efficiency by reducing computing resources, training time, and data collection overhead while achieving high accuracy for specific prediction tasks across different entities in the communication system.
Smart Images

Figure 2025542564000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a communication system and a method for training a machine learning model for a communication system. [Background technology]
[0002] In communication systems, such as 5G mobile communication systems, it is important to ensure that a certain quality of service can be maintained. To this end, various information, such as load, resource usage, available components, user mobility, component status, etc., may be monitored and taken into account when controlling the communication system, e.g., taking measures when overload is imminent to avoid degradation of service quality. The collection and / or evaluation of this information may be performed using machine learning models, which need to be trained appropriately for the corresponding prediction task. Because machine learning models may have different objectives in such contexts, may be used in the context of equipment from various vendors, may relate to different coverage areas, etc., it may be desirable to have multiple machine learning models to be able to have an appropriate machine learning model for all cases. Thus, a high training effort is required to achieve high prediction accuracy for all these cases. Therefore, a technique is desired that enables efficient training of machine learning models for use in communication systems for different objectives, coverage areas, device types (e.g., vendors), etc. Summary of the Invention
[0003] According to one embodiment, 1. A communication system comprising: - a plurality of communication system entities (or (functional) components), each associated with a hierarchical level of a hierarchy of the communication system, with entities associated with higher hierarchical levels of the hierarchy being configured to provide their respective functions to a larger set of subscribers of the communication system than entities associated with lower hierarchical levels of the hierarchy; A training component, obtaining a specification of a first machine learning model trained for a first prediction task for use by a first one of the entities; using the first machine learning model as a basis to generate and train a second machine learning model for a second prediction task for use by a second one of the entities, the second entity being configured to be at a lower hierarchical level than the first entity; a training component; A communication system is provided that includes:
[0004] According to a further embodiment, there is provided a method for training a machine learning model for a communication system in accordance with the above-mentioned communication system.
[0005] In the drawings, like reference numbers generally refer to the same parts throughout the different views. The drawings are not necessarily to scale, emphasis instead being placed upon illustrating the principles of the invention. In the following description, various aspects are described with reference to the following drawings: [Brief explanation of the drawings]
[0006] [Figure 1] FIG. 1 illustrates a communication system according to one embodiment. [Figure 2] FIG. 1 illustrates a communication system with components arranged into levels to show the hierarchy as it exists in a typical (cellular) communication system. [Figure 3] FIG. 1 illustrates the transmission of data for training from a lower level to a higher level. [Figure 4] FIG. 1 illustrates the use of a Machine Learning (ML) Model Repository Function to register information about trained models in a communication system. [Figure 5] FIG. 1 is a flow diagram illustrating model training (e.g., RAN transfer learning procedure) according to one embodiment. [Figure 6] FIG. 10 is a flow diagram illustrating training of an ML model for handover optimization in a base station (e.g., a DU and / or CU) according to one embodiment. [Figure 7] FIG. 10 is a diagram showing an example including five communication hierarchy levels including an AMF (Access and Mobility Management Function) level between an OAM (Operation, Administration and Maintenance) level and a DU level. [Figure 8] , ,Figure 7 illustrates the transport of a UE-specific model in case of a handover of a UE at the second, third, or fourth level of the hierarchy of Figure 7, depending on whether the handover changes only the AMF serving the UE, the CU serving the UE, or the DU serving the UE. [Figure 9] FIG. 1 illustrates an implementation of transfer learning in a communication system having an O-RAN architecture. [Figure 10] FIG. 1 is a flow diagram illustrating a method for training a machine learning model for a communication system. DETAILED DESCRIPTION OF THE INVENTION
[0007] The following detailed description refers to the accompanying drawings, which show, by way of example, specific details and aspects of the present disclosure in which the invention may be practiced. Other aspects may be utilized, and structural, logical, and electrical changes may be made, without departing from the scope of the present invention. Various aspects of the present disclosure are not necessarily mutually exclusive, as some aspects of the present disclosure may be combined with one or more other aspects of the present disclosure to form new aspects.
[0008] Various examples corresponding to aspects of the present disclosure are described below.
[0009] Example 1 is a communications system comprising: a plurality of communications system entities, each associated with a hierarchical level of a hierarchy of the communications system, and entities associated with higher hierarchical levels of the hierarchy configured to provide respective functions to a larger set of subscribers of the communications system than entities associated with lower hierarchical levels of the hierarchy; a training component configured to obtain specifications of a first machine learning model trained for a first predictive task for use by a first one of the entities, and to use the first machine learning model as a basis to generate and train a second machine learning model for a second predictive task for use by a second one of the entities, the second entity being at a lower hierarchical level than the first entity.
[0010] Example 2 is the communication system of Example 1, wherein the first prediction task is a prediction task used by a first entity to provide a first function to a first set of subscribers, and the second prediction task is a prediction task used by a second entity to provide a second function to a second set of subscribers.
[0011] Example 3 is the communication system of example 2, wherein the first prediction task is a prediction task for a first set of subscribers and the second prediction task is the same prediction task for a second set of subscribers.
[0012] Example 4 is the communication system of example 2 or 3, wherein the first set of subscribers includes a second set of subscribers.
[0013] Example 5 is the communication system of example 4, wherein the first set of subscribers includes more subscribers than the second set of subscribers.
[0014] Example 6 is the communication system of any one of Examples 1-5, wherein using the first machine learning model as a basis for generating and training the second machine learning model includes using transfer learning from the first machine learning model to the second machine learning model.
[0015] Example 7 is the communication system of any one of Examples 1 to 6, wherein using the first machine learning model as a basis for generating and training the second machine learning model includes setting initial values of parameters of the second machine learning model to values of corresponding parameters of the first machine learning model.
[0016] An eighth embodiment is the communication system according to any one of the first to seventh embodiments, in which the first machine learning model and the second machine learning model are neural networks, and the parameters are neural network weights.
[0017] Example 9 is the communication system of any one of Examples 1 to 8, wherein using the first machine learning model as a basis for generating and training the second machine learning model includes reusing at least a portion of the first machine learning model for the second machine learning model.
[0018] Example 10 is the communication system of example 9, wherein the first machine learning model and the second machine learning model are neural networks, and the portions are one or more neural network layers.
[0019] Example 11 is the communication system of any one of Examples 1-10, wherein the training component is part of the second entity.
[0020] Example 12 is the communication system of any one of Examples 1-10, wherein the training component is part of an entity different from the second one of the entities.
[0021] Example 13 is the communication system of any one of Examples 1 to 12, wherein the hierarchy includes a plurality of hierarchical levels, and for each hierarchical level of the plurality of hierarchical levels, each entity of the hierarchical level is configured to provide its functionality to serve subscribers within a respective coverage area of the communication system, the size of the coverage area increasing from a higher hierarchical level to a lower hierarchical level of the hierarchy.
[0022] Example 14 is the communications system of any one of Examples 1 to 13, wherein for each respective entity of the entities, the respective entity is configured to provide a respective function to a respective set of subscribers, the entities include a plurality of entities, each of the plurality of entities is configured to provide a respective function to a respective subset of a plurality of subsets into which the respective set of subscribers is divided, and the plurality of entities are at a lower hierarchical level than the respective entity.
[0023] Example 15 is the communication system of any one of Examples 1-14, wherein the entity includes one or more of a plurality of core network functions, a plurality of radio access network components, and a plurality of mobile terminals.
[0024] Example 16 is the communication system of example 15, wherein the plurality of mobile terminals includes a respective mobile terminal for each of the subscribers.
[0025] Example 17 is the communication system of any one of Examples 1 to 16, wherein the entities include one or more of an Operation, Administration and Maintenance system, an Access and Mobility Management Function, a plurality of base stations, a plurality of centralized units, and a plurality of distributed units.
[0026] Example 18 is the communication system of any one of Examples 1 to 17, wherein the second prediction task is a prediction task for a mobile terminal of the communication system, and wherein one of the entities configured to provide a respective function for the mobile terminal using the second machine learning model is configured, in the event of a handover of the mobile terminal, to transfer the second machine learning model to another one of the entities configured to provide a respective function for the mobile terminal after the handover.
[0027] Example 19 is the communication system of any one of Examples 1-18, including a model repository facility configured to store information about trained machine learning models, and wherein the training component is configured to retrieve information about the first machine learning model from the model repository facility and use the retrieved information to obtain a specification of the first machine learning model.
[0028] Example 20 is the communications system of any one of Examples 1-19, wherein the training component is configured to determine base model characteristics from the requirements of the second machine learning model, search the model repository function for a machine learning model having the determined characteristics, and use the machine learning model having the determined characteristics as the first machine learning model.
[0029] Example 21 is the communication system of any one of Examples 1-20, wherein the training component is configured to obtain data for training the second machine learning model from one or more third entities among the entities at a lower hierarchical level than the second entity, and train the second machine learning model using the obtained data.
[0030] Example 22 is the communication system of Example 21, wherein the training component is configured to send rules for filtering and / or aggregating data to one or more third entities, and the one or more third entities are configured to provide data to the second entity for training the machine learning model by filtering and / or aggregating the data according to the rules.
[0031] Example 23 is the communication system of any one of Examples 1 to 22, wherein the second entity is configured to use the trained second machine learning model to perform a second prediction task and to provide a respective function according to an outcome of the second prediction task.
[0032] Example 24 is the communication system of any one of Examples 1 to 23, wherein the first prediction task and / or the second prediction task is a location prediction, a channel condition prediction, a direction prediction, an energy consumption prediction, a measurement value prediction, a trajectory prediction, a communication resource condition prediction, a communication resource usage prediction, a handover prediction, or a data traffic prediction.
[0033] Example 25 is a method of training a machine learning model for a communication system, comprising: determining a hierarchy of the communication system, wherein each communication system entity of a plurality of communication system entities is associated with a hierarchical level of the hierarchy, with entities associated with higher hierarchical levels of the hierarchy being configured to provide functionality to a larger set of subscribers of the communication system than entities associated with lower hierarchical levels of the hierarchy; obtaining specifications of a first machine learning model trained for a first prediction task for use by a first one of the entities; and using the first machine learning model as a basis to generate and train a second machine learning model for a second prediction task for use by a second one of the entities, the second entity being at a lower hierarchical level than the first entity.
[0034] It should be noted that one or more features of any of the above examples may be combined with any one of the other examples, and in particular, embodiments described in the context of an apparatus are equally valid for a method.
[0035] According to further embodiments, there is provided a computer program and computer readable medium comprising instructions which, when executed by a computer, cause the computer to perform the method of any one of the above examples.
[0036] Various examples are described in more detail below.
[0037] FIG. 1 illustrates a communication system according to one embodiment.
[0038] The communication system in this example comprises a Radio Access Network (RAN) 101, a core network 103 and a transport (communication) network 102 connecting the RAN 101 to the core network 103. Subscriber terminals (denoted UE according to 3GPP) 104 are connected to distributed units 105 of the RAN 101 which are connected to a centralized unit 106 (for implementing a base station of the RAN 101). Application functions 107 are connected to the core network 103. The core network 103, for example, a 5G core network (5 GC), includes various network functions such as an AMF (Access and Mobility Management Function) 108, an SMF (Session Management Function) 109, an NEF (Network Exposure Function) 110, an UPF (User Plane Function) 111, an NWDAF (Network Data Analytics Function) 112, and a PCF (Policy Control Function) 113.
[0039] The core network 103 is coupled to an OAM (Operation, Administration and Maintenance) system 114 .
[0040] In communication systems such as that shown in FIG. 1, machine learning (ML) models may be used to perform predictive tasks for a variety of purposes, such as: Enhanced CSI (Channel State Information) feedback, e.g., reducing overhead and improving accuracy of predictions Spatial frequency domain CSI compression using two-sided AI model Time-domain CSI prediction using UE-side model • Beam management, e.g., beam prediction in time and / or spatial domains to reduce overhead and latency and improve beam selection accuracy. Spatial domain downlink beam prediction for the first beam set A based on the measurement results of the second beam set Time downlink beam prediction for the first set of beams based on historical measurements of the second set of beams. To optimize the performance of the radio link between the subscriber terminal 104 and the RAN 101, e.g., one with heavy NLOS (non-line-of-sight) conditions; ·Direct AI / ML positioning AI / ML assisted positioning Improved positioning accuracy for different scenarios, including: Furthermore, AI / ML technologies, e.g. ●Network energy saving, for example, ML models Input: UE mobility / trajectory, current energy efficiency, UE measurement reports Output: predicted energy efficiency, handover strategy Used in Load balancing, for example for ML models, Input: UE trajectory, UE traffic, RAN resource state Output: Predicted resource status information Used in ● Mobility optimization, for example, ML models Input: UE location information, radio measurements, UE handover Output: Predicted handover, UE traffic prediction and may be used, among other things, to improve RAN performance, for example.
[0041] For one of the above purposes, a corresponding ML model may be trained by a single entity (e.g., a base station (denoted as gNB in 5G), OAM 114, AF 107), which collects training data and trains the model, and the model may then be used at multiple inference locations, i.e., by the same entity or other entities (components such as network functions, but also components within the UE). Thus, the training location and the inference location may generally be different.
[0042] To achieve the above objectives, it is necessary to have multiple inference locations (e.g., within the RAN 101). To keep the training effort low, an approach here may be to try to train a common model for all inference locations (e.g., in the OAM 114). However, communication systems typically have a high heterogeneity, i.e., there are different categories of entities in the system with different characteristics (UEs by vendor A, gNBs by vendor B, ...). For example, Regarding CSI use cases: UE from vendor A may use a different method than UE from vendor B. ● Regarding beam management use cases: The beamforming characteristics by a gNB from vendor A may differ from the characteristics of the beam formed by a gNB from vendor B. Regarding positioning use cases: Mobility patterns and positioning techniques in area A may be different from area B. ● Regarding energy saving use cases: The energy consumption model of a gNB from vendor A may be different from a gNB from vendor B.
[0043] Therefore, a common model for all categories needs to be a complex and large model, which is difficult to train, requires a huge amount of training data (to be able to capture differences between multiple categories), and may not achieve high accuracy in all categories. Using a separate model for each category makes it easier to achieve high accuracy, but typically requires fairly lengthy training because models for different categories have much in common. For example, a CSI model for vendor A UE (trained at a first base station) may have much in common with a CSI model for vendor B (trained at a second base station).
[0044] Similarly, ML models for different use cases may have commonalities, e.g., "network energy saving" and "load balancing" use cases may both utilize a UE mobility prediction model, "positioning" and "mobility optimization" use cases may both use a UE mobility prediction model, and "CSI" and "beam management" may both use a channel status prediction model. Thus, while training an individual model for each use case may be a waste of resources, a common model for many use cases (such as for multiple categories) may also be impractical due to high complexity, training data requirements, training effort, and resulting low accuracy.
[0045] According to various embodiments, training techniques are provided that enable training of individual models (e.g., per use case and / or per category) with low requirements on computing resources, training time, and training data collection overhead (compared to the techniques described above), particularly by applying transfer learning across different levels of a communication system hierarchy, i.e., by using multi-level hierarchical transfer learning.
[0046] To this end, it should be noted that communication systems are typically hierarchical in the sense that higher layer components are responsible for larger sets of mobile terminals (e.g., larger geographic areas) and lower layer components are responsible for smaller sets of mobile terminals (e.g., smaller geographic areas).
[0047] FIG. 2 shows a diagram of a communication system 200, where components are arranged at levels 201, 202, 203, 204 to show the hierarchy as it exists in a typical (cellular) communication system.
[0048] As described with reference to FIG. 1, the communication system 200 includes a UE 205 (on the lowest level 204), a DU 206 (on the second lowest level 203), a CU 207 (on the second highest level 202), and an OAM 208 (on the highest level 201).
[0049] The hierarchy arises from the fact that entities associated with higher hierarchical levels in the hierarchy are configured to provide functionality for serving a larger set of subscribers (e.g., serving a larger coverage area) than entities associated with lower hierarchical levels (e.g., serving a smaller coverage area). For example, the OAM 208 is responsible for all subscribers using the communication system, but each DU 206 (as there are multiple of them) is configured to serve only a subset of the subscribers (e.g., subscribers within its associated coverage area). At the lowest level, it can be seen that each UE "serves" (i.e., provides its functionality) only its associated subscribers.
[0050] The UE 205, DU 206, and CU 207 may have different vendors (i.e., be of different types, i.e., implemented differently, and therefore, for example, have different capabilities or performance, etc.), as indicated by different hatching.
[0051] At higher levels, training data for a larger set of UEs or coverage areas or system portions is used to provide a more comprehensive model, which can be transferred to lower levels to build (i.e., generate) more specific models. Lower levels may then build more customized models by adding customization layers to the base model received from the higher levels. Referring to the example of FIG. 2 , the highest level of the OAM is provided with a machine learning model (neural network) 209 having only two layers, while lower levels add layers down to the lowest layer, where the UE-specific neural network 210 has five layers. For any model portions (submodels, such as neural network layers) included in the base model at the higher layer that the entity (at the lower layer) uses to build its individual model, the entity at the lower layer may use the parameters of the base model (e.g., neural network weights) as a starting point for training its individual model (i.e., use transfer learning). Note that individual models may also have different inputs than the base model. For example, additional inputs may be bypassed to layers added to the base model for the individual model.
[0052] Thus, instead of training a common model or training individual models from scratch, a base model may be trained at a higher level and then customized at lower levels (possibly refining the model across different levels). For example, the OAM 308 trains a generic model for "handover optimization" and transfers it to the DU 306 or CU 307 of the gNB, and each gNB (DU or CU) customizes it based on the deployment scenario.
[0053] Each level 201-205 is provided with one or more repositories 211 for storing ML models trained on different entities.
[0054] Not all of the detailed data at lower levels is needed to build (and especially train) models at higher levels; aggregated and / or filtered data sent from lower levels to higher levels may be used.
[0055] FIG. 3 shows the transmission of data for training from a lower level to a higher level.
[0056] As shown in FIG. 2, there are OAM 308, UE 305, DU 306, and CU 307 associated with four hierarchical levels.
[0057] The UE 305 collects detailed UE-specific data 309. The DU 306 collects data from multiple UEs 305. This may include filtering and aggregation (such as averaging across UEs) to produce partially filtered / aggregated data 310. This process continues with the collection of generic aggregated data 311 at the OAM 308. Thus, each level has its own training data that can be used to train machine learning models at each level.
[0058] Filtering and / or aggregation rules that are applied at a level to forward data to the next higher level may be passed down from the higher level. For example, entities at higher levels can subscribe to get data from lower levels. These subscription requests can specify aggregation / filtering rules so that entities at lower levels know how to preprocess the data before sending it to higher levels.
[0059] Therefore, training of base models at higher levels can be done via aggregated information from lower levels (e.g., average number of UEs, average speed, etc.), and this aggregation can be used to reduce the amount of data that needs to be transferred from lower levels to higher levels, thus reducing the training data collection overhead.
[0060] Each model that is trained can be assigned a class to allow lower level components to query and find appropriate models that have been trained at higher levels.
[0061] FIG. 4 illustrates the use of an ML Model Repository Function 410 to register information about trained models in a communication system 400.
[0062] The different entities 405-408 register the models they have trained in the ML model repository function 410 together with the corresponding information (e.g., the model's capabilities, the task for which it was trained, the hierarchical level at which / information about that hierarchical level the model was trained, further information about the training, ...). When other entities 405-408 need an ML model, they can search the repository based on the desired characteristics. Thus, for example, they can obtain the location (e.g., the respective repository 409) of the trained ML model (i.e., its specifications, e.g., regarding the neural network weights) and acquire it from there.
[0063] To enable the above functionality, for example, the following interfaces are provided: A model registration and discovery interface between the ML Model Repository Function (MLRF) 410 and the entities 405-408, which includes, for example, at least the following APIs (Application Programming Interfaces): Model Registration: Entities 405-408 can register ML models and information about them, such as corresponding capabilities and task classes, with the MLRF 410. Model discovery: Entities 405-408 can query the MLRF 410 for information (e.g., location or possibly complete information) about models with given characteristics (e.g., capabilities and tasks for which they are trained). A model distribution interface between higher-level entities and lower-level entities, including at least the following APIs: Model Request: An entity (ML model consumer entity) can request a specific model from another entity (ML model provider entity). Model distribution: The ML model provider entity distributes the requested ML model to the ML model consumer entity. ● A data collection interface between higher level entities and lower level entities, including at least the following APIs: Data Subscription: Higher level entities can subscribe to data available at lower level entities, providing aggregation and / or filtering rules. · Data Provision: Lower level entities provide the requested data according to the rules provided in the subscription.
[0064] Note that the MLRF 410 may be part of the OAM 408.
[0065] FIG. 5 shows a flow diagram 500 illustrating model training (eg, RAN transfer learning procedure) according to one embodiment.
[0066] As described with reference to Figures 2 to 4, the flow involves a first communication system entity 501 at level n-1 (i.e., a lower level entity), a second communication system entity 502 at level n (i.e., a next higher level entity), an MLRF 503, and entities 504 at even higher levels (n+1, n+2, etc.).
[0067] At 505, model training (or retraining) is triggered in the second entity 102, for example because the second entity 502 is to perform a particular task, such as energy management.
[0068] At 506, the second entity 502 identifies the base model required for the model being trained (i.e., identifies the characteristics of the base model, such as the task for which it will be trained).
[0069] At 507, the second entity 502 queries the MLRF 503 to find information about the base model.
[0070] At 508, if the second entity 502 (based on the information obtained from the MLRF 503) finds that it is not already using the base ML model or that there is a newer version of it, it retrieves the model from the source entity.
[0071] At 509, the second entity 502 collects local training data.
[0072] To this end, data is collected from a first (lower level) entity 501 at 510 .
[0073] At 511, a second entity 502 trains a desired model using local training data.
[0074] FIG. 6 shows a flow diagram 600 illustrating training of an ML model for handover optimization in a base station (eg, a DU and / or a CU), according to one embodiment.
[0075] One or more UEs 601, base stations (gNBs, i.e., DUs and / or CUs) 602, MLRFs 603, and OAMs 604 are involved in the flow.
[0076] At 605, handover optimization model training (or retraining) is triggered at gNB 602.
[0077] At 606, the gNB 602 determines that the UE mobility prediction model can be used as the base model.
[0078] At 607, the gNB 602 queries the MLRF 603 for information regarding the UE mobility prediction model.
[0079] At 608, if the gNB 602 has not previously retrieved the UE mobility prediction model, or if the information provided by the MLRF 603 indicates that the UE mobility prediction model has been recently updated, the gNB obtains the latest version from its location, in this example, the OAM 604.
[0080] At 609, the gNB 602 collects or updates local training data.
[0081] To this end, at 610, the gNB 602 collects aggregated and / or filtered data related to handover optimization from one or more UEs 601.
[0082] At 611, the gNB 602 uses the UE mobility prediction model as a base model and the collected training data to train a desired handover optimization model.
[0083] Note that the inference (i.e., model use) can be performed at a variety of locations. 1) Inference at the training location (e.g., UE-specific inference at the same UE): An entity builds a "specific" ML model using a base model (e.g., a general-purpose common model), e.g., adds one or more neural network layers, trains it, and uses the trained model for inference. 2) Inference at level (n+1) entities relative to a particular level n entity (e.g., a gNB makes inferences regarding a particular UE). Level n entities register specific models with the MLRF, level (n+1) entities discover specific models from the MLRF, level (n+1) entities extract models from level n entities, and level (n+1) entities perform inference.
[0084] It is further noted that there are four hierarchical levels (or three if the CU and DU levels are merged into the gNB level, as in the example of Figure 6), but there could be more levels.
[0085] FIG. 7 shows an example including five communication hierarchy levels 701 to 705, including an AMF level 702 between an OAM level 701 and a DU level 703.
[0086] In case of handover, the model can be transferred at different levels as shown in FIG.
[0087] FIG. 8 shows a UE-specific model of transport in case of a handover of a UE 801 on the second, third, or fourth level of the hierarchy of FIG. 7, depending on whether the handover changes only the AMF serving the UE, the CU serving the UE, or the DU serving the UE.
[0088] FIG. 9 illustrates an implementation of transfer learning in a communication system 900 having an O-RAN architecture.
[0089] As in the above example, the communications system 900 has multiple hierarchical levels 901-904 with which various O-RAN architecture components are associated.
[0090] In summary, according to various embodiments, a communications system (e.g., corresponding to communications system 200) is provided that includes a plurality of communications system entities (e.g., corresponding to entities 205-208, but possibly corresponding to other or additional entities), each associated with a hierarchical level (e.g., corresponding to hierarchical levels 201-204, but possibly corresponding to other or additional entities) of the communications system hierarchy, with entities associated with higher hierarchical levels of the hierarchy being configured to provide their respective functions to a larger set of communications system subscribers than entities associated with lower hierarchical levels of the hierarchy. The communication system further includes a training component (e.g., corresponding to one of the entities, such as entities 205-208) configured to obtain specifications of a first machine learning model (e.g., corresponding to ML model 209) trained for a first predictive task for use by a first one of the entities, and to use the first machine learning model as a basis to generate and train a second machine learning model (e.g., corresponding to ML model 210) for a second predictive task for use by a second one of the entities, the second entity being at a lower hierarchical level than the first entity.
[0091] In other words, according to various embodiments, the training efficiency of an ML model (e.g., an NG-RAN ML model) is improved in a communication system, e.g., a 5G communication system, particularly, e.g., a 5GC, by utilizing the hierarchical structure of the mobile communication system using transfer learning. Thus, for example, the efficiency of NG-RAN ML model training for 5GC is improved by using multi-level hierarchical transfer learning. This enables better performance of ML model training, ML model training with less data, less time, and fewer resources. Data collection overhead can be managed through hierarchical multi-level aggregation / filtering.
[0092] It should be noted that the structure of the ML models may be different, for example, the second ML model may include a bypass or one or more additional model heads depending on the inputs and outputs of the first and second ML models.
[0093] Various embodiments provide hierarchical model training and transfer (higher levels provide generic common base models to lower levels), multi-level transfer learning, e.g., in radio access networks, the possibility to extend (further) extended models across multiple levels, and hierarchical data collection (training data can be filtered / aggregated from lower layers to higher layers).
[0094] For example, at level n of a hierarchy (e.g., a gNB), a model is trained (e.g., by the gNB) that is also used for inference at that level n. Instead of redundant training at multiple entities at the same level (e.g., a gNB) on at least partially common features, a common base model (e.g., a common set of neural network layers) may be trained at level n+1 and shared among entities at level n. Even if retraining is necessary at a level, it requires less data and is faster than training a model from scratch at each entity at level n.
[0095] Another possibility is training at level n (e.g., in the gNB) and inference at level n-1 (e.g., in the UE). Thus, instead of training a common model for all (n-1)-level entities, which may lead to accuracy issues, a common base model is trained at level n and customized at level n-1. Furthermore, redundant training is avoided because the base model is trained only once at level n, as opposed to training specific (individual models) for each category of (n-1)-level entities from scratch at level n-1, which may lead to redundant training. By sending aggregated / filtered data from level n-1 to level n, data collection overhead can be kept low.
[0096] Various tasks for improving communication system performance, particularly, for example, RAN performance, may be performed by one or more ML models (e.g., the first ML model and the second ML model of FIG. 9), according to various embodiments.
[0097] For example, a method such as that shown in FIG. 10 is performed.
[0098] FIG. 10 shows a flow diagram 1000 illustrating a method for training a machine learning model for a communication system.
[0099] At 1001, a hierarchy of a communication system is determined (e.g., the communication system is analyzed to determine the hierarchy), wherein each entity of a plurality of communication system entities of the communication system is associated with a hierarchical level of the hierarchy, and wherein entities associated with higher hierarchical levels of the hierarchy are configured to provide functionality to a larger set of subscribers of the communication system than entities associated with lower hierarchical levels of the hierarchy.
[0100] At 1002, a specification of a first machine learning model trained for a first prediction task for use by a first one of the entities is obtained (by a training component).
[0101] At 1003, a second machine learning model is generated and trained using the first machine learning model as a basis for a second prediction task for use by a second one of the entities, the second entity being at a lower hierarchical level than the first entity.
[0102] Components of a communication system (e.g., training components and entities) may be implemented, for example, by one or more circuits. A "circuit" may be understood as any kind of logic implementation entity, and may be a dedicated circuit or a processor that executes software, firmware, or any combination thereof stored in memory. Thus, a "circuit" may be a hard-wired logic circuit or a programmable processor, e.g., a programmable logic circuit such as a microprocessor. A "circuit" may also be software, e.g., a processor that executes any kind of computer program. Any other kind of implementation of each of the above-mentioned functions may also be understood as a "circuit."
[0103] While particular embodiments have been described, it should be understood by those skilled in the art that various changes in form and details can be made therein without departing from the spirit and scope of the embodiments of the present disclosure as defined by the appended claims. The scope is therefore indicated by the appended claims, and all changes that come within the meaning and range of equivalents of the claims are therefore intended to be embraced.
Claims
1. 1. A communication system comprising: a plurality of communication system entities, each associated with a hierarchical level of a hierarchy of the communication system, with entities associated with higher hierarchical levels of the hierarchy being configured to provide respective functions to a larger set of subscribers of the communication system than entities associated with lower hierarchical levels of the hierarchy; A training component, obtaining a specification of a first neural network trained for a first prediction task for use by a first one of the entities; generating a second neural network for a second prediction task for use by a second one of the entities using the first neural network as a basis by adding one or more neural network layers to the first neural network, and training the second neural network, the second entity being configured to be at a lower hierarchical level than the first entity; a training component; Including, Communication system.
2. 2. The communication system of claim 1, wherein the first prediction task is a prediction task used by the first entity to provide a first function to a first set of subscribers, and the second prediction task is a prediction task used by the second entity to provide a second function to a second set of subscribers.
3. 3. The communication system of claim 2, wherein the first prediction task is a prediction task for the first set of subscribers and the second prediction task is the same prediction task for the second set of subscribers.
4. 10. The communication system of claim 1, wherein using the first neural network as a basis for generating and training the second neural network comprises using transfer learning from the first neural network to the second neural network.
5. 2. The communication system of claim 1, wherein using the first neural network as a basis for generating and training the second neural network includes setting initial values of parameters of the second neural network to values of corresponding parameters of the first neural network.
6. 2. The communication system of claim 1, wherein using the first neural network as a basis for generating and training the second neural network includes reusing at least a portion of the first neural network for the second neural network.
7. 2. The communication system of claim 1, wherein the hierarchy includes a plurality of hierarchical levels, and for each hierarchical level of the plurality of hierarchical levels, each entity of the hierarchical level is configured to provide its functionality to serve subscribers within a respective coverage area of the communication system, the size of the coverage area increasing from a higher hierarchical level to a lower hierarchical level of the hierarchy.
8. 2. The communications system of claim 1, wherein for each respective one of said entities, said respective entity is configured to provide a respective function to a respective set of said subscribers, said entities including a plurality of entities, each of said plurality of entities configured to provide a respective function to a respective subset of a plurality of subsets into which said respective set of subscribers is divided, said plurality of entities being at a lower hierarchical level than said respective entity.
9. The communication system of claim 1 , wherein the entities include one or more of a plurality of core network functions, a plurality of radio access network components, and a plurality of mobile terminals.
10. 2. The communication system of claim 1, wherein the second prediction task is a prediction task for a mobile terminal of the communication system, and one of the entities configured to provide a respective function for the mobile terminal using the second neural network is configured to transfer the second neural network to another one of the entities configured to provide a respective function for the mobile terminal after the handover, in the event of a handover of the mobile terminal.
11. 10. The communication system of claim 1, further comprising a model repository facility configured to store information about trained neural networks, wherein the training component is configured to retrieve information about the first neural network from the model repository facility and use the retrieved information to obtain the specification of the first neural network.
12. 2. The communication system of claim 1, wherein the training component is configured to determine base model characteristics from the requirements of the second neural network, search a model repository facility for a neural network having the determined characteristics, and use the neural network having the determined characteristics as the first neural network.
13. 2. The communication system of claim 1, wherein the training component is configured to obtain data for training the second neural network from one or more third entities among entities at a lower hierarchical level than the second entity, and to train the second neural network using the obtained data.
14. 14. The communication system of claim 13, wherein the training component is configured to send rules for filtering and / or aggregating data to the one or more third entities, and the one or more third entities are configured to provide the data to the second entity for training the neural network by filtering and / or aggregating data according to the rules.
15. 1. A method for training a neural network for a communication system, comprising: determining a hierarchy of the communication system, wherein each entity of a plurality of communication system entities of the communication system is associated with a hierarchical level of the hierarchy, and wherein entities associated with higher hierarchical levels of the hierarchy are configured to provide functionality to a larger set of subscribers of the communication system than entities associated with lower hierarchical levels of the hierarchy; obtaining a specification of a first neural network trained for a first prediction task for use by a first one of the entities; generating and training a second neural network for a second prediction task for use by a second one of the entities using the first neural network as a basis by adding one or more neural network layers to the first neural network, the second entity being at a lower hierarchical level than the first entity; A method comprising:
Citation Information
Patent Citations
Transmission device, reception device, transmission method, and reception method
JP2022033051A
Device and method for controlling arrangement of functions of base station and computer program
JP2022165661A
Model management device and model management method
JP2023043082A
Federated learning across UE and ran
US20220038349A1