Distributed learning resource optimization system and method based on open wireless access network

By optimizing resource allocation decisions and introducing an adaptive retransmission mechanism in open radio access networks, the problems of low communication efficiency and degraded model convergence performance in distributed learning are solved, and efficient and reliable distributed learning in open radio access networks is realized.

CN121865296APending Publication Date: 2026-04-14BEIJING JIAOTONG UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING JIAOTONG UNIV
Filing Date
2025-12-11
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing distributed learning methods in 6G networks suffer from problems such as low communication efficiency, increased training latency, decreased model convergence performance, and opaque communication protocols during model parameter uploading. In particular, in scenarios with dense deployment of open radio access networks, channel fading and interference lead to increased training latency and decreased model convergence performance.

Method used

A distributed learning resource optimization system based on an open wireless access network is adopted. The distributed learning resource optimization model is solved by the control unit. Combined with an adaptive retransmission mechanism and a hierarchical aggregation strategy, the resource allocation decision is optimized, including the joint optimization of target user selection, maximum retransmission count, wireless resource blocks, power and computing resources. The association between user groups and access units is established, and the collaborative management of non-real-time intelligent controllers and near-real-time intelligent controllers is introduced to construct an intelligent distributed learning architecture. A two-stage algorithm is used for resource allocation optimization.

Benefits of technology

It improves the reliability and training accuracy of distributed learning, reduces training latency, enhances model convergence performance and communication efficiency, solves the problems of increased training latency and decreased model convergence performance caused by channel fading and interference, and achieves efficient training under low latency and energy consumption constraints.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121865296A_ABST
    Figure CN121865296A_ABST
Patent Text Reader

Abstract

The invention discloses a distributed learning resource optimization system and method based on an open wireless access network, and relates to the technical field of distributed learning, the system comprises a plurality of access units, a concentration unit and a control unit, the access units are in one-to-one correspondence with user groups, the control unit solves a distributed learning resource optimization model, and the concentration unit is in one-to-one correspondence with the user groups. The access unit issues the maximum number of retransmission times, wireless resource blocks, power and computing resources to target users in the user group for the target users to carry out local training and local gradient uploading, and the access unit carries out local aggregation on the local gradients of all the target users in the user group to obtain a local aggregation gradient of the user group; and the concentration unit globally aggregates the local aggregation gradients of all the user groups to obtain a global aggregation gradient. According to the invention, the training time delay can be reduced, the model convergence performance is improved, and the reliability and training precision of distributed learning are fully improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of distributed learning technology, and in particular to a distributed learning resource optimization system and method based on open wireless access networks. Background Technology

[0002] With the rapid development of 6G technology, IoT devices are ubiquitous, generating massive amounts of data for training machine learning models to support applications such as autonomous driving, smart healthcare, and smart manufacturing. Traditional centralized learning methods require transmitting large amounts of data collected from edge devices to the cloud for model training. However, this centralized approach is prone to privacy breaches, is limited by spectrum resources, struggles to support large-scale data transmission, and incurs high communication overhead. Therefore, a learning method that both protects privacy and reduces communication overhead is urgently needed in 6G networks.

[0003] Distributed learning methods enable collaborative model training among multiple users. Users train their models locally using large amounts of collected data, then upload the model parameters to a central server. The central server aggregates the parameters and distributes them back to the users. This process is repeated in multiple rounds of communication until the model converges. While distributed learning methods avoid uploading raw data and, compared to centralized methods, protect privacy and reduce communication overhead, they suffer from low communication efficiency during parameter uploading due to the wide geographical distribution of users. Furthermore, the opaque communication protocol between users and the central server significantly increases the difficulty of troubleshooting and handling transmission errors, hindering timely error processing, increasing training latency, and impacting model performance.

[0004] To address the shortcomings of existing distributed learning methods, Open Radio Access Network (O-RAN) is introduced. O-RAN decomposes the radio access network into Open Radio Units (O-RUs), Open Distributed Units (O-DUs), and Open Centralized Units (O-CUs), achieving hardware and software decoupling and interface standardization. This allows for multi-vendor deployment and supports distributed deployment of edge servers. Users can upload model parameters to the edge servers, which then upload them to the central server, improving model parameter synchronization efficiency and resolving the low communication efficiency issue during model parameter upload in existing distributed learning methods. The communication protocols between Open Radio Units, Open Distributed Units, and Open Centralized Units are transparent, addressing the increased training latency and its impact on model performance in existing distributed learning methods. Furthermore, O-RAN introduces a dual-timescale intelligent controller: a non-real-time intelligent controller and a near-real-time intelligent controller. The non-real-time intelligent controller is responsible for long-term policy generation, while the near-real-time intelligent controller dynamically optimizes network behavior, including retransmission frequency selection and resource allocation, within a timescale of 10 milliseconds to 1 second, thereby supporting efficient training of distributed learning methods.

[0005] However, current distributed learning based on open radio access networks mainly includes two schemes: one is a centralized edge-global aggregation approach, where users first upload their local model parameters to the edge server, the edge server completes the initial aggregation, and then transmits the model parameters obtained from the initial aggregation to the central server for secondary aggregation, and then distributes the global model parameters obtained from the secondary aggregation to the users. However, in scenarios with dense deployment of open radio units, channel fading and interference will lead to increased training latency and decreased model convergence performance. At the same time, the non-independent and identically distributed problem of local datasets of different users will also lead to decreased model convergence performance. The other approach is to introduce a non-real-time intelligent controller for resource scheduling, but under the requirements of low training latency and the constraints of spectrum and energy consumption, there is still a lack of joint optimization strategies for retransmission mechanisms, user scheduling, and multi-dimensional resources, so it is difficult to fully improve the reliability and training accuracy of distributed learning. Summary of the Invention

[0006] The purpose of this application is to provide a distributed learning resource optimization system and method based on open radio access networks, which can reduce training latency, improve model convergence performance, and significantly enhance the reliability and training accuracy of distributed learning.

[0007] To achieve the above objectives, this application provides the following solution.

[0008] In a first aspect, this application provides a distributed learning resource optimization system based on an open radio access network. The distributed learning resource optimization system based on an open radio access network includes: multiple access units, a central unit, and a control unit. Each access unit includes an open radio unit and an open distributed unit of the open radio access network, and the central unit includes an open central unit of the open radio access network. The access unit corresponds one-to-one with the user group, and the access unit is communicatively connected to the user group, the central unit and the control unit respectively. The user group includes several users. The control unit is used to solve the distributed learning resource optimization model to obtain the resource allocation decision for the current communication round. The distributed learning resource optimization model takes minimizing training loss and training latency as the objective function and resource allocation decision value constraints, training latency constraints, and energy consumption constraints as constraints. The resource allocation decision includes the target user and the maximum number of retransmissions, radio resource blocks, power, and computing resources for each target user. The target user is the user in the user group. The access unit is used to send the maximum number of retransmissions, radio resource blocks, power and computing resources to the target users in the user group based on the resource allocation decision of the current communication round. The target user is used to train a machine learning model locally using a local dataset based on the computing resources and the global aggregated gradient of the previous communication round issued by the access unit, to obtain the local gradient of the current communication round, and upload the local gradient of the current communication round to the access unit based on the maximum retransmission count, the radio resource block and the power. The access unit is used to perform local aggregation of the local gradient of the current communication round of all the target users in the user group to obtain the local aggregated gradient of the current communication round of the user group; The centralized unit is used to perform global aggregation of the local aggregated gradients of the current communication rounds of all the user groups to obtain the global aggregated gradient of the current communication round.

[0009] Secondly, this application provides a method for optimizing distributed learning resources based on open radio access networks. Based on the aforementioned system for optimizing distributed learning resources based on open radio access networks, the method includes: In each communication round during the training process: The control unit solves the distributed learning resource optimization model to obtain the resource allocation decision for the current communication round. The distributed learning resource optimization model takes minimizing training loss and training latency as the objective function and resource allocation decision value constraints, training latency constraints, and energy consumption constraints as constraints. The resource allocation decision includes the target user and the maximum number of retransmissions, radio resource blocks, power, and computing resources for each target user. Based on the resource allocation decision of the current communication round, the access unit sends the maximum retransmission count, radio resource blocks, power, and computing resources to the target users in the user group. The target users use the computing resources and the global aggregated gradient of the previous communication round sent by the access unit to train the machine learning model locally using the local dataset to obtain the local gradient of the current communication round, and upload the local gradient of the current communication round to the access unit based on the maximum retransmission count, the radio resource blocks, and the power. The access unit performs local aggregation on the local gradient of the current communication round of all the target users in the user group to obtain the local aggregated gradient of the current communication round of the user group; The centralized unit performs global aggregation of the local aggregated gradients of the current communication rounds for all the user groups to obtain the global aggregated gradient of the current communication round.

[0010] According to the specific embodiments provided in this application, this application has the following technical effects: This application provides a distributed learning resource optimization system and method based on an open radio access network. The control unit solves the distributed learning resource optimization model to obtain the resource allocation decision for the current communication round. The distributed learning resource optimization model takes minimizing training loss and training latency as the objective function, and resource allocation decision value constraints, training latency constraints, and energy consumption constraints as constraints. The resource allocation decision includes the target user and the maximum number of retransmissions for each target user, radio resource blocks, power, and computing resources. Thus, under the constraints of low training latency requirements (objective function) and spectrum (resource allocation decision value constraints and training latency constraints) and energy consumption (energy consumption constraints), the retransmission mechanism (maximum number of retransmissions), user scheduling (target user), and multi-dimensional resources (radio resource blocks, power, and computing resources) can be jointly optimized, which can significantly improve the reliability and training accuracy of distributed learning. When a target user completes local training and uploads the local gradient of the current communication round to the access unit, retransmission can be performed under the constraint of the maximum number of retransmissions. By introducing a retransmission mechanism and determining the optimal maximum number of retransmissions, the upload success rate can be improved, reducing the increase in training latency and the decrease in model convergence performance caused by channel fading and interference. This reduces training latency, improves communication efficiency, and enhances model convergence performance. An access unit is configured to correspond one-to-one with a user group. The access unit performs local aggregation of the local gradients of all target users in the user group for the current communication round, obtaining the locally aggregated gradient for the current communication round of the user group. The central unit then performs global aggregation of the locally aggregated gradients of all user groups for the current communication round, obtaining the globally aggregated gradient for the current communication round. This reduces the model convergence performance degradation caused by the non-independent and identically distributed nature of the local datasets of different users, thus improving model convergence performance. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 The flowchart illustrates the implementation of the distributed learning resource optimization system based on an open wireless access network provided in Embodiment 1 of this application.

[0013] Figure 2 This is a diagram of the intelligent distributed learning architecture provided in Embodiment 1 of this application.

[0014] Figure 3 This is a diagram of the distributed learning management architecture provided in Embodiment 1 of this application.

[0015] Figure 4A flowchart illustrating the optimization of model accuracy and training latency provided in Embodiment 1 of this application.

[0016] Figure 5 The flowchart illustrates the two-stage algorithm provided in Embodiment 1 of this application for optimizing model accuracy and training latency.

[0017] Figure 6 This is a flowchart illustrating a distributed learning resource optimization method based on an open wireless access network, as provided in Embodiment 2 of this application. Detailed Implementation

[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0019] Example 1.

[0020] This embodiment provides a distributed learning resource optimization system based on an open radio access network. The distributed learning resource optimization system based on an open radio access network includes: multiple access units, a central unit, and a control unit. Each access unit includes an open radio unit and an open distributed unit of the open radio access network, and the central unit includes an open central unit of the open radio access network.

[0021] Each access unit corresponds one-to-one with a user group. The access unit is connected to the user group, the central unit, and the control unit for communication. Each user group includes several users.

[0022] The control unit is used to solve the distributed learning resource optimization model to obtain the resource allocation decision for the current communication round. The distributed learning resource optimization model takes minimizing training loss and training latency as the objective function, and resource allocation decision value constraints, training latency constraints, and energy consumption constraints as constraints. The resource allocation decision includes the target user and the maximum number of retransmissions for each target user, radio resource blocks, power, and computing resources. The target user is the user in the user group.

[0023] The access unit is used to make resource allocation decisions based on the current communication round, and to send the maximum number of retransmissions, radio resource blocks, power and computing resources to the target users in the user group.

[0024] The target user uses the global aggregated gradient of the previous communication round issued by the access unit based on computing resources and the local dataset to train the machine learning model locally, obtain the local gradient of the current communication round, and upload the local gradient of the current communication round to the access unit based on the maximum number of retransmissions, radio resource blocks and power.

[0025] The access unit is used to perform local aggregation of the local gradients of the current communication rounds of all target users in the user group, so as to obtain the local aggregated gradient of the current communication rounds of the user group.

[0026] The centralized unit is used to perform global aggregation of the local aggregated gradients of all user groups in the current communication round, so as to obtain the global aggregated gradient of the current communication round.

[0027] This embodiment aims to provide a distributed learning resource optimization system based on open radio access networks (RANs) to address the problems of low communication efficiency, data non-independent and identically distributed characteristics affecting model accuracy, and complex resource scheduling in RAN-based distributed learning. To address the low communication efficiency, an adaptive retransmission mechanism is employed to improve the success rate of data (local gradient) uploads. To address the issue of data non-independent and identically distributed characteristics affecting model accuracy, a two-layer aggregation strategy (local aggregation at access units and global aggregation at centralized units) is designed, and an association between user groups and access units is established (i.e., one user group corresponds to one access unit), improving the convergence performance of the global model (i.e., the machine learning model with parameters determined by the global aggregation gradient after training). To address the complex resource scheduling issue, a distributed learning resource optimization model is established to optimize model accuracy and training latency. This model is solved using a two-stage algorithm to optimize model accuracy and training latency, achieving joint optimization of user selection (determining target users), maximum retransmission count, radio resource blocks, power, and computing resources. This enables resource scheduling for RAN-based distributed learning, significantly improving training performance and enhancing system robustness, thereby significantly improving the reliability and training accuracy of distributed learning.

[0028] This embodiment can be applied to distributed learning scenarios, especially in complex and dynamic wireless network environments. Figure 1The flowchart illustrates the implementation of a distributed learning resource optimization system based on an open radio access network. First, an intelligent distributed learning architecture consisting of users, access units, and centralized units is constructed to lay the foundation for subsequent model training. An adaptive retransmission mechanism is introduced into this architecture to ensure communication stability under different network environments and improve the data upload success rate during model training. Next, collaborative management of non-real-time and near-real-time intelligent controllers is introduced to construct a distributed learning management architecture that coordinates these controllers. This enables centralized management and intelligent control of the open radio access network during distributed learning, ensuring closed-loop optimization of the distributed learning process to optimize resource allocation and improve system responsiveness. Then, a joint optimization mathematical model for model accuracy and training latency in intelligent distributed learning is established, resulting in a distributed learning resource optimization model. This achieves a balance between model accuracy and training latency, thereby accelerating the convergence speed of the global model. Finally, a two-stage algorithm is designed to solve the distributed learning resource optimization model, completing resource allocation during the distributed learning process and addressing the resource allocation problem. This achieves collaborative optimization of user selection, maximum retransmission count, radio resource blocks, power, and computing resources to improve model training efficiency and convergence performance.

[0029] In a first aspect, this embodiment provides an intelligent distributed learning architecture composed of users, access units, and a centralized unit. Figure 2 This diagram illustrates an intelligent distributed learning architecture, primarily composed of users, access units, and centralized units. At the user end, model initialization begins, including gradient uploading, K-means clustering, and user access, establishing the association between user groups and access units. Model training then proceeds, including resource management, gradient reception, and gradient descent, yielding the local gradient for the current communication round. At the access and centralized units, model transmission occurs first, including gradient transmission, confirmation, and gradient retransmission. The local gradient for the current communication round is uploaded to the access unit. Subsequently, model aggregation is performed, including edge aggregation (local aggregation at the access unit), global aggregation (global aggregation at the centralized unit), and gradient broadcasting, distributing the globally aggregated gradient for the current communication round to users for the next communication round. This entire process achieves closed-loop management from user local training and gradient uploading to global aggregation and distribution at the centralized unit, thus supporting adaptive retransmission and hierarchical optimization in distributed learning.

[0030] By employing an adaptive retransmission mechanism and a hierarchical distributed learning framework, communication efficiency and task execution performance in the distributed learning process can be effectively improved. Subsequently, a two-stage algorithm is introduced, combining a non-real-time intelligent controller and a near-real-time intelligent controller, to dynamically adjust the user's retransmission strategy and the allocation strategies for wireless resource blocks and computing resources. This solves the model generalization problem caused by communication interruptions and the non-independent and identically distributed data distribution among different users, improving the model's accuracy while reducing training latency. The key is to construct an adaptive retransmission mechanism that can optimize communication efficiency and resource allocation in the distributed learning process. This embodiment introduces this adaptive retransmission mechanism into the intelligent distributed learning architecture. The intelligent distributed learning architecture will be described in detail below.

[0031] To address the issue of data loss during data uploads in existing distributed learning systems on open radio access networks, an intelligent distributed learning architecture is proposed, such as... Figure 2 As shown, the intelligent distributed learning architecture of this embodiment consists of users, access units, and centralized units. The access unit is an integrated architecture of open radio units and open distributed units, including initialization process, resource management process, user model training process, user gradient transmission and retransmission process, access unit and centralized unit aggregation process, and model broadcasting and update process.

[0032] The initialization process takes place in a distributed learning scenario within an open wireless access network. The centralized unit initializes model parameters, specifically initializing global model parameters based on task requirements. These initial parameters are then distributed to all users. For each user, the access unit closest to them is initially designated as their corresponding access unit. The centralized unit transmits the initial model parameters to the access units, which in turn distribute them to the users corresponding to those access units. Upon receiving the initial model parameters, users train their machine learning model using their local dataset (local, non-independent, identically distributed data) to obtain initial local gradients (the gradient is the derivative of the loss function with respect to the model parameters, used to determine these parameters). These initial local gradients are then uploaded to the user's corresponding access unit (the access unit closest to the user), which in turn transmits them back to the centralized unit. Compared to uploading model parameters, uploading local gradients offers stronger privacy protection and higher communication efficiency. After collecting the initial local gradients of all users, the centralized unit uses the K-means clustering algorithm to group all users based on gradient similarity (cosine similarity can be used to characterize the similarity of users' initial local gradients), resulting in multiple user groups. For each user group, the average distance from each user in the user group to the same access unit is calculated to obtain the average distance corresponding to the access unit. The above operation is performed for each access unit to obtain the average distance corresponding to each access unit. The access unit with the shortest (i.e., the smallest) average distance is selected as the access unit corresponding to the user group, thus determining the association between user groups and access units, so that one user group corresponds to one access unit. At this time, the user groups and access units are in a one-to-one correspondence, which can alleviate the model bias caused by the non-independent and identically distributed data and improve the model convergence performance.

[0033] The K-means clustering algorithm can also be replaced by other clustering algorithms. When using the K-means clustering algorithm for grouping, K users are first randomly selected as initial cluster centers. For each user, the cosine similarity between the user's initial local gradient and the initial local gradient of each initial cluster center is calculated. The user is assigned to the initial cluster center with the highest cosine similarity, resulting in K initial clusters. The average cosine similarity between the user in each initial cluster and the initial cluster center is calculated to obtain the initial local gradient of the cluster center after the initial cluster is updated. The next iteration is performed until the maximum number of iterations is reached, resulting in K clusters. Users belonging to the same cluster are grouped into K user groups.

[0034] The resource management process begins at the start of each training round when the control unit generates user scheduling and resource allocation strategies (i.e., resource allocation decisions) based on user distribution (i.e., user grouping) and channel conditions. These resource allocation decisions are then distributed to the access units, including user selection, maximum retransmission count, radio resource blocks, power, and computing resources. This enables the joint allocation of user selection, maximum retransmission count, radio resource blocks, power, and computing resources, ensuring optimal utilization under limited resource conditions.

[0035] The user model training process involves the target user (i.e., the user selected through user selection) receiving the global aggregated gradient from the previous communication round, performing local training based on the allocated computing resources, and using the mini-batch stochastic gradient descent algorithm to calculate the local gradient for the current communication round.

[0036] The user gradient transmission and retransmission process involves the target user uploading its local gradient for the current communication round to the access unit using allocated radio resource blocks and power. The access unit's link layer performs a CRC check to determine whether the user's uploaded local gradient has been successfully received and correctly decoded, returning an acknowledgment or non-acknowledgment flag. If the flag is acknowledgment, the transmission is successful; otherwise, if the flag is non-acknowledgment, the transmission has failed, and a retry is initiated, triggering a retransmission. This process continues until the allocated maximum number of retransmissions is reached or the transmission is successful, thus introducing an adaptive retransmission mechanism to improve the reliability of gradient uploading and the robustness of the model.

[0037] The aggregation process between the access unit and the centralized unit involves the access unit performing local aggregation (also known as edge aggregation) on the local gradient of each target user in the current communication round of the received (successfully received) user group, obtaining the local aggregated gradient of the user group in the current communication round, and transmitting the local aggregated gradient to the centralized unit. The centralized unit then performs global aggregation on the local aggregated gradients of all received user groups in the current communication round, obtaining the global aggregated gradient of the current communication round. By introducing an adaptive retransmission mechanism and joint optimization, it can consider users with channel instability and retransmission impact, fully taking into account the impact of wireless channel errors, packet loss, and retransmission, and ensuring the stability of the aggregation result.

[0038] The model broadcasting and update process involves completing local and global aggregation. The centralized unit transmits the global aggregated gradient to the access unit, which then distributes (broadcasts) the global aggregated gradient to all users. Users update their local model parameters based on the received global aggregated gradient, completing the current communication round and waiting for the next communication round of local training. This enables the distributed learning model to converge quickly and efficiently coordinate communication resources in a dynamic communication environment.

[0039] Secondly, this embodiment proposes a distributed learning management architecture for collaborative non-real-time intelligent controllers and near-real-time intelligent controllers. Figure 3This is a diagram of a distributed learning management architecture, which defines a data plane and a control plane, i.e., it is divided into two layers: the data plane and the control plane.

[0040] In the data plane, to achieve low-latency wireless resource control, open radio units and open distributed units are jointly deployed to form access units. Since retransmission-related protocols are mainly implemented in open distributed units, open centralized units act as centralized units, handling higher-level protocols implemented in them. The centralized units are responsible for global model aggregation. Users connect to the access units via the air interface, and the access units connect to the centralized units via the F1 interface for parameter transmission. Communication between the access units and the centralized units follows the standard interface of the wireless access network, ensuring the timeliness and standardization of data exchange between various network components, and realizing collaborative management and parameter synchronization between distributed and centralized units.

[0041] On the control plane, the non-real-time intelligent controller and the near-real-time intelligent controller work together. Specifically, the non-real-time intelligent controller collects user distribution and channel status through the O1 interface to support resource allocation decisions. The non-real-time intelligent controller and the near-real-time intelligent controller formulate strategies based on user distribution and channel status. They exchange strategy information through the A1 interface to achieve coordinated control and obtain resource allocation decisions. The non-real-time intelligent controller transmits the resource allocation decisions to the access unit through the E2 interface for resource allocation.

[0042] Specifically, in order to manage the resource allocation of the distributed learning process, a data plane is designed. The data plane ensures the data transmission and computation in the distributed learning process. On the data plane, the access unit, which is formed by the collaborative deployment of open radio units and open distributed units, can shorten the feedback link and support low-latency adaptive retransmission. The access unit communicates with the centralized unit through the F1 interface and transmits the user gradient data to the centralized unit for global aggregation in a timely manner. Open radio access networks introduce non-real-time intelligent controllers (NRTICs) and near-real-time intelligent controllers (NRTICs) as control planes. On the control plane, the NRTIC periodically collects long-term statistical information such as user distribution, channel status, and computing power. It obtains user distribution and channel status from the access unit through Open Interface 1 (O1) and uses this information to cooperate with the NRTIC in formulating resource allocation decisions. In this process, the NRTIC determines user selection, maximum retransmission count, and radio resource blocks, and transmits the strategy (user selection, maximum retransmission count, and radio resource blocks) to the NRTIC through the A1 interface. The NRTIC then determines power and computing resources and transmits the strategy (power and computing resources) back to the NRTIC through the A1 interface. The NRTIC combines user selection, maximum retransmission count, radio resource blocks, power, and computing resources to obtain resource allocation decisions, which are then distributed to the access unit through the E2 interface. This achieves joint allocation of user selection, maximum retransmission count, radio resource blocks, power, and computing resources, thereby reducing training latency and energy consumption while ensuring model convergence accuracy. Through the collaborative design of the control plane and data plane, the system can perceive the network status in real time in a dynamic wireless environment and intelligently adjust the resource allocation strategy.

[0043] Thirdly, this embodiment provides a mathematical model for the joint optimization of model accuracy and training latency in intelligent distributed learning, which is a mathematical expression of the problem of optimizing model accuracy and training latency. Figure 4 The flowchart for optimizing model accuracy and training latency includes communication and computation models for user retransmissions, as well as modeling a multi-objective optimization problem considering multi-dimensional resource constraints. Specifically, firstly, a distributed learning communication model is constructed, including uplink transmission of the local model and downlink distribution of edge and global models. Secondly, a distributed learning computation model is constructed to describe the computation process of the local model. Finally, based on the above distributed learning communication and distributed learning computation models, a collaborative optimization problem for model accuracy and training latency (i.e., a multi-objective optimization problem considering multi-dimensional resource constraints) is proposed. User selection, maximum retransmission count, radio resource blocks, power, and computational resources are jointly optimized to achieve a balance between communication efficiency and learning performance.

[0044] The communication and computation models for user retransmission (i.e., the distributed learning communication model and the distributed learning computation model) refer to the following: the transmission of distributed learning parameters mainly includes two parts: uplink transmission of the local model and downlink transmission of the edge and global models. Therefore, transmission latency and energy consumption need to be calculated. Since fiber optic transmission has higher reliability and lower latency than wireless channels, the global aggregation transmission latency and packet errors between the access unit and the central unit can be ignored. In the distributed learning computation process, compared to local training, the aggregation latency and energy consumption of the edge and global models are negligible; therefore, only the local computation process is considered. Specifically, the training metrics for distributed learning parameters mainly include local training latency and local training energy consumption.

[0045] Among them, modeling a multi-objective optimization problem that considers multi-dimensional resource constraints means that, under the premise of meeting energy consumption requirements, it is necessary to minimize the training loss and training latency of distributed learning, and jointly decide on user selection, maximum retransmission count, wireless resource blocks, power, and the allocation of computing resources to achieve optimal resource utilization and training effect.

[0046] Specifically, to optimize model accuracy and training latency in distributed learning, it is necessary to establish communication and computation models for user retransmission. Specifically, during edge aggregation, OFDMA (Orthogonal Frequency Division Multiple Access) is considered for uploading local gradients within each access unit, and the channel is modeled as a quasi-static fading channel. Defined as a set of radio resource blocks. To accelerate the uplink process, the local gradient is divided into multiple data packets for transmission. The transmission rate of the data packets is calculated using finite block length theory. Therefore, the access unit... Corresponding user (i.e., access unit) Users in the corresponding user group In communication rounds The transmission rate is: ; in, Access Unit Corresponding user In communication rounds The transmission rate; This represents the number of wireless resource blocks. It is a binary variable. If wireless resource blocks In communication rounds It was assigned to the access unit Corresponding user ,but ,otherwise, ; This refers to the bandwidth of the radio resource block during uplink. The mathematical expectation of the channel gain is given, and its calculation process is a well-established technique that will not be elaborated upon here. Access Unit Corresponding user In communication rounds via wireless resource blocks Channel gain during data packet upload; Access Unit Corresponding user In communication rounds via wireless resource blocks Signal-to-interference-plus-noise ratio (SIR / NDR) during data packet upload; This refers to the data packet size. It is the inverse function of the Gaussian Q-function; For users The probability of decoding errors.

[0047] This embodiment assumes that each user can be allocated at most one radio resource block. and access unit Signal-to-interference-plus-noise ratio for: ; in, Access Unit Corresponding user In communication rounds The power; For other users in communication rounds Use wireless resource blocks The interference caused specifically refers to other users during communication rounds. Use wireless resource blocks Interference caused by transmitting other business data (not used for distributed learning) shall be determined by the user based on the actual situation; Let be the power spectral density of the noise.

[0048] Assume that the local gradient magnitude is the same for each user, and that the local gradient is divided into... Data packets, therefore, user Upload latency of the first local gradient upload for: ; in, Access Unit Corresponding user In communication rounds The initial upload delay; The number of data packets obtained from local gradient partitioning.

[0049] Since the expected channel gain is the same for each radio resource block within a communication round, and each data packet is transmitted independently, therefore the user If the transmission time of each data packet is the same, then the user In communication rounds The energy consumption for the first data packet transmission is: ; in, Access Unit Corresponding user In communication rounds The energy consumption of the first upload.

[0050] However, due to interference signals and noise in the wireless environment, data packets may be corrupted during the initial transmission. The access unit employs a cyclic redundancy check (CRC) scheme to check for errors in the data packets. When a user receives a non-acknowledgment from the access unit, the user can retransmit multiple times, retransmitting the user's locally assigned data packets. The error rate of the first transmitted data packet is: ; in, Access Unit Corresponding user In communication rounds The error rate of the first transmitted data packet; This is a preset function. , As variables, For integration variables; The length of the block; For the channel dispersion function, ; For bitrate, .

[0051] user The total packet error rate in the first transmission was , It is a linear normalization function, based on the packet error rate. The upper and lower bounds are obtained, and the value of the linear normalization function is equal to (packet error rate - lower bound of packet error rate) / (upper bound of packet error rate - lower bound of packet error rate). The number of data packets obtained from local gradient partitioning.

[0052] Because the user experiences independent random noise and interference, they assume each retransmission is independent; therefore, the user... Local gradient retransmission Indicator function to indicate whether the success was successful for: ; Among them, when When the time is right, it means the retransmission was successful; otherwise, it means the retransmission was unsuccessful. When the time is up, it means the retransmission failed. In mathematical terms, the probability refers to the probability of at least one successful transmission. Access Unit Corresponding user In communication rounds The maximum number of retransmissions.

[0053] At the start of each communication round, the access unit or central unit broadcasts the global aggregation gradient to the users. Downlink bandwidth is significantly greater than uplink bandwidth, therefore all radio resource blocks can be used. Downlink transmission time is a fixed value compared to uplink transmission time. Since the access unit and the central unit have power supply access, the energy consumption of transmission can be ignored.

[0054] In distributed learning computation, the aggregation latency and energy consumption of edge and global models are negligible compared to local training. Therefore, considering only the local computation process, the main training metrics include local training latency and local training energy consumption.

[0055] Each user performs within one communication round. The local iteration, in order to simplify the model, uses the CPU for the local iteration process. Therefore, the user... The local training latency is: ; in, Access Unit Corresponding user In communication rounds Local training latency; This represents the number of local iterations. For users Calculate the CPU cycles required for a sample dataset; The number of sample data points that need to be computed in one local iteration; Access Unit Corresponding user In communication rounds The computing resources, specifically the user CPU computing power (CPU cycles per second).

[0056] Since a user's energy is limited, the uplink process is constrained by the energy consumed by the computing resources used by the user. Therefore, the local training energy consumption for each user is: ; in, Access Unit Corresponding user In communication rounds Local training energy consumption; For users The equivalent capacitance coefficient depends on the chip architecture.

[0057] Considering the uplink communication and computation processes, in one communication round The training latency is: ; in, For communication rounds Training latency; For the target user set.

[0058] This embodiment considers a synchronous distributed learning process, so the total training latency depends on the user with the longest training latency.

[0059] Energy consumption per user: ; in, Access Unit Corresponding user In communication rounds Energy consumption.

[0060] To optimize model accuracy and training latency in distributed learning, decisions regarding user selection, maximum retransmission count, radio resource blocks, power, and computational resources need to be made collaboratively, while adhering to energy and resource constraints. Therefore, the optimization problem can be defined as: ; in, This represents the number of communication rounds from the first communication round to the current communication round. For communication rounds The training loss is equal to the number of communication rounds. The weighted sum of the local losses of the target users, where the local loss is the loss value calculated using the loss function. For loss function, For communication rounds The model parameters can be obtained from the convergence analysis of distributed learning. Upper boundary; This is the time delay weighting factor.

[0061] After setting the range of values ​​for the decision variables, training latency, and energy consumption constraints, the above optimization problem can be solved.

[0062] The resource allocation decision value constraints are as follows: ; ; ; ; ; ; in, Total number of users; This represents the maximum number of retransmissions. This is the upper limit of power; This is the lower limit for computing resources; To calculate the upper limit of resources.

[0063] The training latency constraint is: ; in, This is the upper limit for training latency.

[0064] Energy consumption constraints are: ; in, For users The upper limit of energy consumption.

[0065] In the fourth aspect, this embodiment provides a two-stage algorithm to solve the distributed learning resource optimization model and address the resource allocation problem in intelligent distributed learning. Figure 5 The flowchart illustrates a two-stage algorithm for optimizing model accuracy and training latency, demonstrating how the optimization problems of training latency and model accuracy are decomposed into two sub-problems and solved separately. In the first stage, power control and computational resource allocation are modeled as convex optimization problems, and the optimal solution is found using the Lagrange dual method, achieving optimal power and computational resource allocation. In the second stage, user selection, maximum retransmission count selection, and radio resource block allocation are modeled as non-convex problems, and reinforcement learning is used for intelligent decision optimization to achieve optimal user selection, maximum retransmission count, and radio resource block allocation. Regarding algorithm deployment, the first-stage algorithm is deployed on a near-real-time intelligent controller, while the second-stage algorithm is deployed on a non-real-time intelligent controller, achieving hierarchical collaborative optimization and improving the communication efficiency of distributed learning.

[0066] The decomposition into two subproblems is derived from the discrete nature of the decision variables in the original optimization problem. Based on the convergence upper bound of distributed learning, the objective function is approximated as a closed-form problem, and then decomposed into two subproblems according to the discrete nature of the decision variables. Subproblem one is the joint power and computational resource allocation subproblem. Given decisions regarding user selection, maximum retransmission count, and radio resource blocks, the original optimization problem can be reformulated as subproblem one. Subproblem two is the user selection, maximum retransmission count, and radio resource block allocation subproblem. Given decisions regarding power and computational resources, the original optimization problem can be reformulated as subproblem two.

[0067] For the first-stage algorithm (i.e., the stage one algorithm) for solving subproblem one, the original non-convex optimization problem is first transformed by variable substitution. Replace the variable form with ( The newly introduced intermediate variable is used as a variable, thus transforming the original non-convex optimization problem into a convex optimization problem, which can be defined as: ; in, The target number of users.

[0068] The equality constraint is: ; The first term of the convex optimization problem is obtained based on the convergence analysis of distributed learning. The upper realm.

[0069] Subsequently, the Lagrangian function for the convex optimization problem was constructed. It can be written as: ; in, It is the first Lagrange multiplier; It is the second Lagrange multiplier.

[0070] Next, using duality theory, the Lagrange function is equivalently transformed into its corresponding dual problem, the expression of which is: ; The constraints are: ; ; The first Lagrange multiplier can be obtained by solving. Second Lagrange multiplier The value of .

[0071] Finally, the Kuhn-Tuck (KKT) conditions are solved, i.e., further iterated using the KKT conditions, to obtain the optimal solutions for the two decision variables of power and computational resources, which can be written as: ; The power and computational resources can be calculated.

[0072] For the second-stage algorithm (i.e., the stage two algorithm) for solving subproblem two, since the decision variables in subproblem two are discrete, it is a non-convex optimization problem. Traditional algorithms suffer from high computational complexity and can only obtain approximate solutions, failing to effectively handle situations with a large number of wireless resource blocks and requiring long-term decision optimization. Therefore, this embodiment proposes a resource allocation algorithm based on deep reinforcement learning. This algorithm models subproblem two as a Markov decision process. Through interaction between the agent and the network environment, it learns the optimal decision variables, thereby optimizing user selection, maximum retransmission count, and wireless resource block allocation. At each decision moment, the agent observes the current network state and selects an action. Subsequently, based on environmental feedback rewards and state transitions, it enters the next network state. Through continuous iterative learning, it obtains the optimal action, thus optimizing the action decision-making process.

[0073] In reinforcement learning, the state space, action space, and reward are defined as follows.

[0074] In the state space, the non-real-time intelligent controller needs to collect the network state in each communication round. Since the optimization objective is to minimize training loss and training latency, the network state should include both direct and indirect metrics affecting the objective function. Specifically, the network state is designed to include channel gain, training latency, energy consumption, packet loss rate, and constants related to hyperparameters in the user's previous communication round, which can be described as: ; in, For communication rounds Network status; , , users respectively In communication rounds The training latency, energy consumption, and packet loss rate, For communication rounds The constant term related to hyperparameters in the loss function; For the total set of users.

[0075] In the action space, when the network state is observed, the non-real-time intelligent controller makes a temporary resource allocation decision, resulting in an action. According to sub-problem two, the user's action is divided into two parts, and the action is defined as follows: ; in, For communication rounds The action. Target user set. It can be determined by the allocated radio resource blocks, i.e. .

[0076] In the reward, when the non-real-time intelligent controller performs an action It will receive rewards from the environment. To obtain the optimal user selection, maximum retransmission count, and radio resource block, the reward includes training loss, training latency, and a penalty term. The reward function can be defined as: ; in, For communication rounds The reward; For the target user set; It is the first weighting factor; For communication rounds The training loss is equal to the number of communication rounds. The weighted sum of the local losses of the target users, For loss function, For communication rounds Model parameters; This is the time delay weighting factor; For communication rounds Training latency; As the second weighting factor; Access Unit Corresponding user In communication rounds The training latency is equal to ; This is the upper limit for training latency; For a binary indicator function, when When the inequality within the parentheses is true, If it is 0, otherwise, =-1; It is the third weighting factor; Access Unit Corresponding user In communication rounds Energy consumption; For users The upper limit of energy consumption.

[0077] In the reward function, the first term is the reward for model accuracy and training latency obtained from the upper bound of distributed learning convergence; the second and third terms are the penalties for violating training latency and energy consumption limits. By maximizing the cumulative reward, the optimal resource allocation decision is obtained.

[0078] The computational complexity of the Phase 1 algorithm is significantly lower than that of the Phase 2 algorithm. Therefore, the Phase 1 algorithm is suitable for execution on near real-time intelligent controllers, while the Phase 2 algorithm is suitable for operation on non-real-time intelligent controllers. Furthermore, by using the Phase 1 algorithm as an inner sub-algorithm of the Phase 2 algorithm, it only needs to be called once in each optimization round during the reinforcement learning model training process. Compared to the alternating optimization method, this significantly accelerates the overall convergence speed. Specifically, within the Phase 2 algorithm, the Phase 1 algorithm can be used to solve for the optimal power and computational resource allocation for any user given a maximum number of retransmissions and a radio resource block.

[0079] This embodiment proposes an intelligent and efficient distributed learning resource allocation method by introducing an adaptive retransmission mechanism, a hierarchical aggregation strategy, and a two-stage algorithm. This method effectively improves the communication efficiency and resource scheduling performance of distributed learning in open radio access networks. The method employs an adaptive retransmission mechanism to optimize the data transmission process and reduce the risk of communication interruption. By designing a two-layer aggregation strategy and the association between users and access units, it addresses the problem of decreased model convergence performance caused by non-independent and identically distributed data. Simultaneously, the two-stage algorithm jointly optimizes user selection, maximum retransmission count, radio resource blocks, power, and computational resources based on task characteristics and network conditions, significantly improving model accuracy and reducing training latency.

[0080] This embodiment details the intelligent distributed learning architecture, the distributed learning management architecture, the mathematical formulation of joint optimization of model accuracy and training latency, and the two-stage algorithm combining convex optimization and deep reinforcement learning. Based on this, the distributed learning resource optimization system based on open radio access networks provided in this embodiment includes: multiple access units, a central unit, and a control unit. Each access unit includes an open radio unit and an open distributed unit of the open radio access network, and the central unit includes an open central unit of the open radio access network.

[0081] Each access unit corresponds one-to-one with a user group. The access unit is connected to the user group, the central unit, and the control unit for communication. Each user group includes several users.

[0082] The process of determining the one-to-one correspondence between access units and user groups includes: for each user, the access unit closest to the user is initially defined as the access unit corresponding to the user; the centralized unit is also used for initialization, generating initial model parameters; the access unit is also used to distribute the initial model parameters to the user corresponding to the access unit; the user is also used to train the machine learning model using the local dataset based on the initial model parameters to obtain the initial local gradient; the access unit is also used to upload the initial local gradient to the centralized unit; the centralized unit is also used to group all users using the K-means clustering algorithm based on the initial local gradients of all users, obtaining multiple user groups, and selecting the access unit with the closest average distance as the access unit corresponding to the user group, so that the access unit and the user group correspond one-to-one, and the average distance is the average distance from each user in the user group to the access unit.

[0083] The control unit is used to solve the distributed learning resource optimization model to obtain the resource allocation decision for the current communication round. The distributed learning resource optimization model takes minimizing training loss and training latency as the objective function and resource allocation decision value constraints, training latency constraints, and energy consumption constraints as constraints. The resource allocation decision includes the target user and the maximum number of retransmissions for each target user, radio resource blocks, power, and computing resources. The target user is the user in the user group.

[0084] The control unit includes a non-real-time intelligent controller and a near-real-time intelligent controller, which are communicatively connected. Both the non-real-time and near-real-time intelligent controllers are also communicatively connected to the access unit. In solving the distributed learning resource optimization model to obtain the resource allocation decision for the current communication round, the non-real-time intelligent controller in the control unit is used to acquire the network state. Using the network state as input, it uses a trained reinforcement learning model to determine the target user, maximum retransmission count, and radio resource block for the current communication round. The target user, maximum retransmission count, and radio resource block for the current communication round are then input into the distributed learning resource optimization model to obtain the post-input optimization model. The post-input optimization model is then transformed to obtain a convex optimization model. The network state includes the channel gain of each user when uploading data packets through the radio resource block in the current communication round, as well as the training latency, energy consumption, packet loss rate, and constant terms related to hyperparameters in the loss function for each user in the previous communication round.

[0085] The near real-time intelligent controller in the control unit is used to solve the convex optimization model to obtain the power and computing resources of the current communication round.

[0086] The non-real-time intelligent controller in the control unit is used to combine the target user, maximum retransmission count, radio resource block, power and computing resources of the current communication round to obtain the resource allocation decision for the current communication round.

[0087] The non-real-time intelligent controller in the control unit is also used to train the trained reinforcement learning model.

[0088] The access unit is used to make resource allocation decisions based on the current communication round, and to send the maximum number of retransmissions, radio resource blocks, power and computing resources to the target users in the user group.

[0089] The target user uses the global aggregated gradient of the previous communication round issued by the access unit based on computing resources and the local dataset to train the machine learning model locally, obtain the local gradient of the current communication round, and upload the local gradient of the current communication round to the access unit based on the maximum number of retransmissions, radio resource blocks and power.

[0090] Based on the global aggregated gradient of the previous communication round issued by the computing resources and access unit, the machine learning model is trained locally using the local dataset to obtain the local gradient of the current communication round. Then, based on the maximum retransmission count, radio resource blocks, and power, the local gradient of the current communication round is uploaded to the access unit. The target user uses this gradient for: Based on the global aggregate gradient of the previous communication round issued by the access unit, the model parameters are calculated and then input into the machine learning model to obtain the training model. Based on computing resources, the training model is trained locally using a local dataset to obtain the local gradient for the current communication round. The local gradient of the current communication round is divided into multiple data packets, and the data packets are uploaded to the access unit through radio resource blocks using power. The system retrieves a flag from the access unit indicating whether the upload was successful. If the flag indicates a successful upload, the system uploads the next data packet until all data packets have been uploaded. If the flag indicates a failed upload, the system retransmits the data packet until the maximum number of retransmissions is reached.

[0091] The access unit is used to perform local aggregation of the local gradients of the current communication rounds of all target users in the user group, so as to obtain the local aggregated gradient of the current communication rounds of the user group.

[0092] In terms of local aggregation of the local gradients of all target users in the user group for the current communication round to obtain the local aggregated gradient of the user group for the current communication round, the access unit calculates the proportion of the local dataset of each target user in the user group to the first total data volume to obtain the weight of the target user; based on the weights of all target users in the user group, the local gradients of all target users in the user group for the current communication round are weighted and summed to perform local aggregation to obtain the local aggregated gradient of the user group for the current communication round; the first total data volume is the sum of the local datasets of all target users in the user group.

[0093] The centralized unit is used to perform global aggregation of the local aggregated gradients of all user groups in the current communication round, so as to obtain the global aggregated gradient of the current communication round.

[0094] In terms of globally aggregating the local aggregated gradients of all user groups in the current communication round to obtain the global aggregated gradient of the current communication round, the centralized unit is used to calculate the proportion of the first total data volume of the user group to the second total data volume for each user group, so as to obtain the weight of the user group; based on the weights of all user groups, the local aggregated gradients of all user groups in the current communication round are weighted and summed to perform global aggregation, so as to obtain the global aggregated gradient of the current communication round; the second total data volume is the sum of the first total data volume of all user groups.

[0095] This embodiment provides an application scenario in network traffic prediction. Users train a machine learning model locally using locally collected base station traffic time-series data (including uplink and downlink traffic volume, active user count, and connection count, forming a local dataset). The machine learning model can be an LSTM, using MSE as the loss function to minimize the error between the predicted and actual traffic values. After training through a distributed learning method, the resulting global model can more accurately predict future base station traffic, providing real-time intelligent decision support for open radio access network (RAN) slice allocation and edge computing task placement.

[0096] Example 2.

[0097] This embodiment provides a distributed learning resource optimization method based on open radio access networks, which works based on the distributed learning resource optimization system based on open radio access networks described in Embodiment 1, such as... Figure 6 As shown, the distributed learning resource optimization method based on open wireless access networks includes the following steps.

[0098] During each communication round in the training process, perform the following steps.

[0099] S1, the control unit solves the distributed learning resource optimization model to obtain the resource allocation decision for the current communication round; the distributed learning resource optimization model takes minimizing training loss and training latency as the objective function, and resource allocation decision value constraints, training latency constraints and energy consumption constraints as constraints; the resource allocation decision includes the target user and the maximum number of retransmissions, radio resource blocks, power and computing resources for each target user.

[0100] S2, the access unit, based on the resource allocation decision of the current communication round, sends the maximum retransmission count, radio resource block, power, and computing resources to the target user in the user group; the target user uses the computing resources and the global aggregated gradient of the previous communication round sent by the access unit to train the machine learning model locally using the local dataset to obtain the local gradient of the current communication round, and uploads the local gradient of the current communication round to the access unit based on the maximum retransmission count, the radio resource block, and the power.

[0101] S3, the access unit performs local aggregation on the local gradient of the current communication round of all the target users in the user group to obtain the local aggregated gradient of the current communication round of the user group.

[0102] S4, the centralized unit performs global aggregation of the local aggregated gradients of the current communication rounds of all the user groups to obtain the global aggregated gradient of the current communication round.

[0103] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0104] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A distributed learning resource optimization system based on open wireless access networks, characterized in that, The distributed learning resource optimization system based on open radio access network includes: multiple access units, a central unit, and a control unit. Each access unit includes an open radio unit and an open distributed unit of the open radio access network, and the central unit includes an open central unit of the open radio access network. The access unit corresponds one-to-one with the user group, and the access unit is communicatively connected to the user group, the central unit and the control unit respectively. The user group includes several users. The control unit is used to solve the distributed learning resource optimization model to obtain the resource allocation decision for the current communication round. The distributed learning resource optimization model takes minimizing training loss and training latency as the objective function and resource allocation decision value constraints, training latency constraints, and energy consumption constraints as constraints. The resource allocation decision includes the target user and the maximum number of retransmissions, radio resource blocks, power, and computing resources for each target user. The target user is the user in the user group. The access unit is used to send the maximum number of retransmissions, radio resource blocks, power and computing resources to the target users in the user group based on the resource allocation decision of the current communication round. The target user is used to train a machine learning model locally using a local dataset based on the computing resources and the global aggregated gradient of the previous communication round issued by the access unit, to obtain the local gradient of the current communication round, and upload the local gradient of the current communication round to the access unit based on the maximum retransmission count, the radio resource block and the power. The access unit is used to perform local aggregation of the local gradient of the current communication round of all the target users in the user group to obtain the local aggregated gradient of the current communication round of the user group; The centralized unit is used to perform global aggregation of the local aggregated gradients of the current communication rounds of all the user groups to obtain the global aggregated gradient of the current communication round.

2. The distributed learning resource optimization system based on open wireless access network according to claim 1, characterized in that, For each user, the access unit that is closest to the user is initially defined as the access unit corresponding to the user; The centralized unit is also used for initialization, generating initial model parameters; The access unit is also used to send the initial model parameters to the user corresponding to the access unit; The user is also used to train the machine learning model using a local dataset based on the initial model parameters to obtain the initial local gradient; The access unit is also used to upload the initial local gradient to the central unit; The centralized unit is further configured to group all users based on the initial local gradients of all users using the K-means clustering algorithm to obtain multiple user groups, and select the access unit with the closest average distance as the access unit corresponding to the user group, so that the access unit corresponds one-to-one with the user group; the average distance is the average distance from each user in the user group to the access unit.

3. The distributed learning resource optimization system based on open wireless access network according to claim 1, characterized in that, The objective function is: ; in, This represents the number of communication rounds from the first communication round to the current communication round. For communication rounds The training loss is equal to the number of communication rounds. The weighted sum of the local losses of the target users, For loss function, For communication rounds Model parameters; This is the time delay weighting factor; For communication rounds Training latency; The resource allocation decision value constraint is as follows: ; ; ; ; ; ; in, If it is a binary variable, such as a wireless resource block In communication rounds It was assigned to the access unit Corresponding user ,but ; Total number of users; This represents the number of wireless resource blocks. Access Unit Corresponding user In communication rounds Maximum number of retransmissions; This represents the maximum number of retransmissions. Access Unit Corresponding user In communication rounds The power; This is the upper limit of power; This is the lower limit for computing resources; Access Unit Corresponding user In communication rounds Computing resources; To calculate the upper limit of resources; The training latency constraint is: ; in, This is the upper limit for training latency; The energy consumption constraint is: ; in, Access Unit Corresponding user In communication rounds Energy consumption; For users The upper limit of energy consumption.

4. The distributed learning resource optimization system based on open wireless access network according to claim 3, characterized in that, The formula for calculating the training latency is: ; ; ; ; ; in, For the target user set; Access Unit Corresponding user In communication rounds Local training latency; Access Unit Corresponding user In communication rounds The initial upload delay; This represents the number of local iterations. For users Calculate the CPU cycles required for a sample dataset; The number of sample data points that need to be computed in one local iteration; The number of data packets obtained from local gradient partitioning; This refers to the data packet size. Access Unit Corresponding user In communication rounds The transmission rate; The bandwidth of the wireless resource block; The mathematical expectation of the channel gain. Access Unit Corresponding user In communication rounds via wireless resource blocks Channel gain during data packet upload; Access Unit Corresponding user In communication rounds via wireless resource blocks Signal-to-interference-plus-noise ratio (SIR / NDR) during data packet upload; It is the inverse function of the Gaussian Q-function; For users The probability of decoding errors; For other users in communication rounds Use wireless resource blocks The resulting interference; Let be the power spectral density of the noise.

5. The distributed learning resource optimization system based on open wireless access network according to claim 4, characterized in that, The formula for calculating energy consumption is: ; ; ; in, Access Unit Corresponding user In communication rounds Local training energy consumption; Access Unit Corresponding user In communication rounds Energy consumption during the first upload; For users The equivalent capacitance coefficient.

6. The distributed learning resource optimization system based on open wireless access network according to claim 1, characterized in that, The control unit includes a non-real-time intelligent controller and a near-real-time intelligent controller. The non-real-time intelligent controller and the near-real-time intelligent controller are communicatively connected, and both are communicatively connected to the access unit. In solving the distributed learning resource optimization model to obtain the resource allocation decision for the current communication round, the non-real-time intelligent controller in the control unit acquires the network state. Using the network state as input, it uses a trained reinforcement learning model to determine the target user, maximum retransmission count, and radio resource block for the current communication round. The target user, maximum retransmission count, and radio resource block for the current communication round are then input into the distributed learning resource optimization model to obtain an input-optimized model. This input-optimized model is then transformed to obtain a convex optimization model. The network state includes the channel gain of each user when uploading data packets through the radio resource block in the current communication round, as well as the training latency, energy consumption, packet loss rate, and constant terms related to hyperparameters in the loss function for each user in the previous communication round. The near real-time intelligent controller in the control unit is used to solve the convex optimization model to obtain the power and computing resources of the current communication round; The non-real-time intelligent controller in the control unit is used to combine the target user, maximum retransmission count, radio resource block, power, and computing resources of the current communication round to obtain the resource allocation decision for the current communication round.

7. The distributed learning resource optimization system based on open wireless access network according to claim 6, characterized in that, The non-real-time intelligent controller in the control unit is also used to train a pre-trained reinforcement learning model, wherein the reward function used when training the pre-trained reinforcement learning model is: ; in, For communication rounds The reward; For the target user set; It is the first weighting factor; For communication rounds The training loss is equal to the number of communication rounds. The weighted sum of the local losses of the target users, For loss function, For communication rounds Model parameters; This is the time delay weighting factor; For communication rounds Training latency; As the second weighting factor; Access Unit Corresponding user In communication rounds Training latency; This is the upper limit for training latency; For a binary indicator function, when When the inequality within the parentheses is true, If it is 0, otherwise, =-1; It is the third weighting factor; Access Unit Corresponding user In communication rounds Energy consumption; For users The upper limit of energy consumption.

8. The distributed learning resource optimization system based on open wireless access network according to claim 1, characterized in that, Based on the global aggregated gradient of the previous communication round issued by the computing resources and the access unit, the machine learning model is trained locally using the local dataset to obtain the local gradient of the current communication round. Then, based on the maximum retransmission count, the radio resource block, and the power, the local gradient of the current communication round is uploaded to the access unit. The target user is used for: Based on the global aggregate gradient of the previous communication round issued by the access unit, the model parameters are calculated and input into the machine learning model to obtain the training model. Based on the computing resources, the training model is trained locally using the local dataset to obtain the local gradient for the current communication round. The local gradient of the current communication round is divided into multiple data packets, and the data packets are uploaded to the access unit through the radio resource block using the power. Obtain the flag returned by the access unit indicating whether the upload was successful or not. If the flag indicates that the upload was successful, upload the next data packet until all data packets have been uploaded. If the flag indicates that the upload failed, retransmit the data packet again until the maximum number of retransmissions is reached.

9. The distributed learning resource optimization system based on open wireless access network according to claim 1, characterized in that, In terms of local aggregation of the local gradients of the current communication rounds of all target users in the user group to obtain the local aggregated gradient of the current communication rounds of the user group, the access unit is used to calculate the proportion of the local dataset of each target user in the user group to the first total data volume, and obtain the weight of the target user. Based on the weights of all target users in the user group, the local gradients of the current communication rounds of all target users in the user group are weighted and summed to perform local aggregation, thereby obtaining the local aggregated gradient of the current communication rounds of the user group; the first total data volume is the sum of the data volumes of the local datasets of all target users in the user group; In terms of globally aggregating the local aggregated gradients of all user groups in the current communication round to obtain the global aggregated gradient of the current communication round, the centralized unit is used to calculate, for each user group, the proportion of the first total data volume of the user group to the second total data volume to obtain the weight of the user group; based on the weights of all user groups, the local aggregated gradients of all user groups in the current communication round are weighted and summed to perform global aggregation to obtain the global aggregated gradient of the current communication round; the second total data volume is the sum of the first total data volume of all user groups.

10. A method for optimizing distributed learning resources based on open radio access networks, operating according to the distributed learning resource optimization system based on open radio access networks as described in any one of claims 1-9, characterized in that, The distributed learning resource optimization method based on open radio access networks includes: In each communication round during the training process: The control unit solves the distributed learning resource optimization model to obtain the resource allocation decision for the current communication round. The distributed learning resource optimization model takes minimizing training loss and training latency as the objective function and resource allocation decision value constraints, training latency constraints, and energy consumption constraints as constraints. The resource allocation decision includes the target user and the maximum number of retransmissions, radio resource blocks, power, and computing resources for each target user. Based on the resource allocation decision of the current communication round, the access unit sends the maximum retransmission count, radio resource blocks, power, and computing resources to the target users in the user group. The target users use the computing resources and the global aggregated gradient of the previous communication round sent by the access unit to train the machine learning model locally using the local dataset to obtain the local gradient of the current communication round, and upload the local gradient of the current communication round to the access unit based on the maximum retransmission count, the radio resource blocks, and the power. The access unit performs local aggregation on the local gradient of the current communication round of all the target users in the user group to obtain the local aggregated gradient of the current communication round of the user group; The centralized unit performs global aggregation of the local aggregated gradients of the current communication rounds for all the user groups to obtain the global aggregated gradient of the current communication round.