A computing power management method for distributed AI training tasks in a mobile network
By constructing a mobile private network integrating sensing and computing, the problem of interruption of distributed AI training tasks during mobile network switching was solved, realizing dynamic reallocation and synchronization of computing resources, and ensuring task continuity and computing efficiency.
Patent Information
- Application Number
- CN202510097812.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-01-22
AI Technical Summary
In mobile networks, distributed AI training tasks are prone to interruptions, duplicate data transmissions, and wasted computing resources during switching, making it impossible to effectively guarantee continuity.
A mobile private network integrating sensing, computing, and computing is constructed by adding network elements for network-computer interface fusion sensing, control scheduling, and service orchestration on the core network side. This allows the network and computing resource status to be perceived, AI training tasks to be split into distributed collaborative sub-tasks, and computing resource reallocation and synchronization to be performed during network switching, ensuring task continuity.
It achieves business continuity during the switching process of distributed AI training tasks, reduces redundant data transmission and waste of computing resources, and shortens training time.
Smart Images

Figure CN119922636B_ABST
Abstract
Description
Technical Field
[0001] This invention discloses a computing power management method for distributed AI training tasks in mobile networks, relating to the fields of mobile communication and artificial intelligence technologies. Background Technology
[0002] With the rapid development of mobile communication technology and the widespread application of AI technology, distributed AI training is increasingly used in mobile networks. However, when mobile terminals switch within a mobile network, such as during Xn or N2 switching, the continuity of distributed AI training tasks is often not effectively guaranteed due to the dynamic changes in computing power nodes and network quality, which can easily lead to problems such as training interruption, duplicate data transmission, and waste of computing resources. Summary of the Invention
[0003] This invention addresses the problems of existing technologies by providing a computing power management method for distributed AI training tasks in mobile networks. By enhancing the core network to construct an integrated mobile private network of communication, sensing, and computing, computing power resources are reallocated and the progress of old and new computing power nodes is synchronized during network switching, ensuring the continuity of distributed AI training tasks during the switching process.
[0004] The specific solution proposed in this invention is as follows:
[0005] This invention also provides a method for managing computing power in distributed AI training tasks in mobile networks, comprising:
[0006] Constructing an integrated mobile private network combining sensing, computing, and computing: Adding converged computing and network sensing network elements, converged computing and network control and scheduling network elements, and converged computing and network service orchestration network elements to the core network. The converged computing and network sensing network elements perceive the status of network and computing resources; the converged computing and network control and scheduling network elements control the scheduling strategies for network and computing resources; and the converged computing and network service orchestration network elements orchestrate distributed AI training tasks.
[0007] When a mobile terminal accesses the integrated mobile private network for communication, sensing, and computing, a session is established for the mobile terminal's AI training task through the core network. The network elements, through the converged computing and network service orchestration, break down the AI training task into distributed collaborative training sub-tasks for computation based on the computing resources of the mobile terminal, base station, and cloud node.
[0008] When a mobile terminal triggers a handover, the core network identifies the network handover signaling and performs the corresponding AI training task computing service handover based on the network handover signaling.
[0009] Furthermore, the computing power service switching in the computing power management method for distributed AI training tasks in a mobile network includes:
[0010] The core network identifies the Path Switch Request network handover signaling and performs an XN handover of the mobile terminal's computing power services based on the network handover signaling, resulting in a change of computing power nodes.
[0011] The core network identifies the network handover signaling Handover Required and performs N2 handover of the computing power service of the mobile terminal according to the network handover signaling, causing a change in the computing power node.
[0012] Furthermore, the XN switching in the aforementioned method for managing computing power for distributed AI training tasks in a mobile network, which causes changes in computing power nodes, includes:
[0013] After the core network identifies the network handover signaling Path Switch Request, it switches the user SUPI, target base station Target-RAN-ID, and handover type Xn through the computing-network convergence service orchestration network element. The computing-network convergence service orchestration network element finds the distributed AI training task identifier TASKID based on the user SUPI and updates the computing power nodes and training strategy of the AI training task. When the Xn handover begins, it stops updating parameters with other computing power nodes, notifies the target base station computing power node to reserve computing power resources and warm up the model, and synchronizes the parameters of the old computing power node to the new computing power node through the Xn interface between the source base station and the target base station. After the handover, the new computing power node resumes computing based on the existing training experience of the old computing power node and performs collaborative training synchronization with the cloud node.
[0014] Furthermore, the N2 switching of the computing power management method for distributed AI training tasks in a mobile network, which causes a change in computing power nodes, includes:
[0015] When the core network identifies the HandOver Requested network handover signaling, it switches the user SUPI, target base station Target-RAN-ID, and handover type N2 through the computing-network convergence service orchestration network element. The computing-network convergence service orchestration network element finds the distributed AI training task identifier TASKID based on the user SUPI and updates the computing power nodes and training strategy of the AI training task. During the N2 handover preparation phase, it notifies the target base station computing power nodes to reserve computing power resources and warm up the model. During the handover, it stops updating parameters with other computing power nodes and synchronizes the old computing power node parameters to the new computing power node through the transmission tunnel between the source base station, the core network UPF, and the target base station. After the handover, the new computing power node resumes computing based on the existing training experience of the old computing power node and performs collaborative training synchronization with the cloud node.
[0016] This invention also provides a computing power management device for distributed AI training tasks in a mobile network, including a mobile private network management module and a computing power management module.
[0017] The mobile private network management module constructs an integrated mobile private network combining sensing, computing, and computing: It adds converged computing and network sensing network elements, converged computing and network control and scheduling network elements, and converged computing and network service orchestration network elements to the core network side. The converged computing and network sensing network elements perceive the status of network and computing resources; the converged computing and network control and scheduling network elements control the scheduling strategies for network and computing resources; and the converged computing and network service orchestration network elements orchestrate distributed AI training tasks.
[0018] When a mobile terminal accesses the integrated mobile private network, the core network establishes a session for the mobile terminal's AI training task through the computing power management module. The computing power management module, through network elements that orchestrate computing and network convergence services, breaks down the AI training task into distributed collaborative training sub-tasks for computation based on the computing resources of the mobile terminal, base station, and cloud node.
[0019] When a mobile terminal triggers a handover, the computing power management module identifies the network handover signaling through the core network and executes the corresponding AI training task computing power service handover based on the network handover signaling.
[0020] Furthermore, the computing power management module of the computing power management device for distributed AI training tasks in a mobile network performs computing power service switching, including:
[0021] The core network identifies the Path Switch Request network handover signaling and performs an XN handover of the mobile terminal's computing power services based on the network handover signaling, resulting in a change of computing power nodes.
[0022] The core network identifies the network handover signaling Handover Required and performs N2 handover of the computing power service of the mobile terminal according to the network handover signaling, causing a change in the computing power node.
[0023] Furthermore, the computing power management module of the computing power management device for distributed AI training tasks in a mobile network performs an XN switch, causing a change in computing power nodes, including:
[0024] After the core network identifies the network handover signaling Path Switch Request, it switches the user SUPI, target base station Target-RAN-ID, and handover type Xn through the computing-network convergence service orchestration network element. The computing-network convergence service orchestration network element finds the distributed AI training task identifier TASKID based on the user SUPI and updates the computing power nodes and training strategy of the AI training task. When the Xn handover begins, it stops updating parameters with other computing power nodes, notifies the target base station computing power node to reserve computing power resources and warm up the model, and synchronizes the parameters of the old computing power node to the new computing power node through the Xn interface between the source base station and the target base station. After the handover, the new computing power node resumes computing based on the existing training experience of the old computing power node and performs collaborative training synchronization with the cloud node.
[0025] Furthermore, the computing power management module of the computing power management device for distributed AI training tasks in a mobile network performs an N2 switch, causing a change in computing power nodes, including:
[0026] When the core network identifies the HandOver Requested network handover signaling, it switches the user SUPI, target base station Target-RAN-ID, and handover type N2 through the computing-network convergence service orchestration network element. The computing-network convergence service orchestration network element finds the distributed AI training task identifier TASKID based on the user SUPI and updates the computing power nodes and training strategy of the AI training task. During the N2 handover preparation phase, it notifies the target base station computing power nodes to reserve computing power resources and warm up the model. During the handover, it stops updating parameters with other computing power nodes and synchronizes the old computing power node parameters to the new computing power node through the transmission tunnel between the source base station, the core network UPF, and the target base station. After the handover, the new computing power node resumes computing based on the existing training experience of the old computing power node and performs collaborative training synchronization with the cloud node.
[0027] The advantages of this invention are:
[0028] This invention solves the problem of interruption of distributed AI training services due to terminal mobility in mobile communication networks. It realizes dynamic switching of computing power tasks, ensures the service continuity of AI training tasks, reduces repeated data transmission and calculation, saves computing power resources, and shortens training time. Attached Figure Description
[0029] Figure 1 This is a schematic diagram before the network architecture switchover.
[0030] Figure 2 This is a schematic diagram after the network architecture has been switched.
[0031] Figure 3 This is a schematic diagram of the N2 switching process.
[0032] Figure 4 This is a schematic diagram of the XN switching process. Detailed Implementation
[0033] The integration of communication, sensing, and computing involves 6G networks. It deeply integrates the communication functions of mobile networks with AI computing functions, providing mobile terminals with available computing resources anytime, anywhere.
[0034] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, so that those skilled in the art can better understand and implement the present invention. However, the embodiments described are not intended to limit the present invention.
[0035] This invention is particularly applicable to scenarios where the continuity of distributed AI training tasks cannot be guaranteed due to changes in computing power nodes and network quality when a mobile terminal triggers a switch.
[0036] Example 1
[0037] This invention also provides a method for managing computing power in distributed AI training tasks in mobile networks, comprising:
[0038] Constructing an integrated mobile private network: Adding computing-network converged sensing network elements, computing-network converged control and scheduling network elements, and computing-network converged service orchestration network elements to the core network side. The computing-network converged sensing network elements perceive the status of network and computing resources, the computing-network converged control and scheduling network elements control the scheduling strategies of network and computing resources, and the computing-network converged service orchestration network elements orchestrate distributed AI training tasks. These network elements interact with other programs through API interfaces based on HTTP2 service messages.
[0039] When a mobile terminal accesses the integrated mobile private network for communication, sensing, and computing, a session is established for the mobile terminal's AI training task through the core network. The network elements, through the converged computing and network service orchestration, break down the AI training task into distributed collaborative training sub-tasks for computation based on the computing resources of the mobile terminal, base station, and cloud node.
[0040] When a mobile terminal triggers a handover, the core network identifies the network handover signaling and performs the corresponding AI training task computing service handover based on the network handover signaling.
[0041] The aforementioned computing power service switching may include:
[0042] The core network identifies the Path Switch Request network handover signaling and performs an XN handover of the mobile terminal's computing power services based on the network handover signaling, resulting in a change of computing power nodes.
[0043] The core network identifies the network handover signaling Handover Required and performs N2 handover of the computing power service of the mobile terminal according to the network handover signaling, causing a change in the computing power node.
[0044] Specifically, the N2 switch, which causes a change in computing power nodes, can include:
[0045] The core network detects an impending N2 handover by recognizing the Handover Required network handover signaling and extracts information such as the user's SUPI and the target base station's Target-RAN-ID from the signaling.
[0046] The core network notifies the computing-network converged service orchestration network element that the mobile terminal is preparing for handover. This is done by specifying the user SUPI, target base station Target-RAN-ID, and handover type N2 through the computing-network converged service orchestration network element.
[0047] The computing-network convergence service orchestration network element finds the computing task identifier (TASKID) under the current terminal through the user's SUPI and requests a new resource scheduling policy from the computing-network convergence control and scheduling network element.
[0048] The network elements for computing-network convergence service orchestration orchestrate services according to the new resource scheduling strategy. After notifying the computing nodes on the target base station to start reserving the distributed computing power indicators and pre-loading the operating environment, they quickly test-run the training model.
[0049] Upon receiving the Handover Request Ack signaling, confirming that the mobile terminal can initiate an N2 handover, the core network, after returning the Handover Command signaling, notifies the edge cloud computing nodes to stop updating gradient parameters and transmitting training data to the source base station computing nodes. Simultaneously, it notifies the source base station computing nodes to stop training, record the number of training epochs, gradients, weights, biases, and other parameters, and package the untrained data.
[0050] The source and target base stations synchronize parameters and untrained data through a fronthaul tunnel established with the core network UPF.
[0051] Upon receiving the Handover Notify signaling from the target base station, the core network detects the completion of the handover and informs the computing-network convergence service orchestration network element. This network element then notifies the edge cloud computing nodes to continue updating gradients and transmitting training data to the target base station. Simultaneously, it notifies the source base station computing nodes to clear old task information. After the handover, the new computing nodes resume computation based on the existing training experience of the old nodes and perform collaborative training synchronization with the cloud computing nodes.
[0052] Specifically, XN switching causes changes in computing power nodes, including:
[0053] The core network detects that a mobile terminal is undergoing an Xn handover by identifying the Path Switch Request network handover signaling, and extracts information such as SUPI and Target-RAN-ID from the signaling.
[0054] The core network notifies the mobile terminal of the computing-network converged service orchestration network element that a handover is in progress. The handover process involves specifying the user SUPI, target base station Target-RAN-ID, and handover type XN through the computing-network converged service orchestration network element.
[0055] The computing-network convergence service orchestration network element finds the computing task identifier (TASKID) under the current terminal through the user's SUPI and requests a new resource scheduling policy from the computing-network convergence control and scheduling network element.
[0056] The network element for converged computing and network services orchestrates services according to the new resource scheduling strategy. It immediately notifies the edge cloud computing power nodes to stop updating gradient parameters and transmitting training data to the source base station computing power nodes. At the same time, it notifies the computing power nodes on the target base station to start reserving the distributed computing power indicators and pre-loading the operating environment, as well as quickly trial-running the training model. It also notifies the source base station computing power nodes to stop training and records the number of training rounds, gradient, weight, bias and other parameters, as well as packaged untrained data.
[0057] The source base station and the target base station directly synchronize parameters and untrained data through the Xn interface between the base stations.
[0058] After the core network completes the downlink tunnel update with the target base station, it sends a Path SwitchRequest Ack signaling message to the target base station, indicating that the downlink is operational. Then, the core network notifies the computing-network convergence service orchestration network element that the handover is complete. The computing-network convergence service orchestration network element instructs the edge cloud computing power nodes to begin updating gradients and transmitting data to the target base station, while also instructing the source base station computing power nodes to clear old task information. After the handover, the new computing power nodes resume computation based on the existing training experience of the old computing power nodes and perform collaborative training synchronization with the cloud computing power nodes.
[0059] Example 2
[0060] This invention also provides a computing power management device for distributed AI training tasks in a mobile network, including a mobile private network management module and a computing power management module.
[0061] The mobile private network management module constructs an integrated mobile private network combining sensing, computing, and computing: It adds converged computing and network sensing network elements, converged computing and network control and scheduling network elements, and converged computing and network service orchestration network elements to the core network side. The converged computing and network sensing network elements perceive the status of network and computing resources; the converged computing and network control and scheduling network elements control the scheduling strategies for network and computing resources; and the converged computing and network service orchestration network elements orchestrate distributed AI training tasks.
[0062] When a mobile terminal accesses the integrated mobile private network, the core network establishes a session for the mobile terminal's AI training task through the computing power management module. The computing power management module, through network elements that orchestrate computing and network convergence services, breaks down the AI training task into distributed collaborative training sub-tasks for computation based on the computing resources of the mobile terminal, base station, and cloud node.
[0063] When a mobile terminal triggers a handover, the computing power management module identifies the network handover signaling through the core network and executes the corresponding AI training task computing power service handover based on the network handover signaling.
[0064] The information interaction and execution process between the modules in the above-mentioned device are based on the same concept as the method embodiment of the present invention, and the specific details can be found in the description in the method embodiment of the present invention, and will not be repeated here.
[0065] Similarly, the device of the present invention solves the problem of interruption of distributed AI training services due to terminal mobility in mobile communication networks, realizes dynamic switching of computing power tasks, ensures the service continuity of AI training tasks, reduces repeated data transmission and calculation, saves computing power resources, and reduces training time.
[0066] It should be noted that not all steps and modules in the above processes and device structures are mandatory; some steps or modules can be omitted as needed. The execution order of each step is not fixed and can be adjusted as required. The system structure described in the above embodiments can be a physical structure or a logical structure. That is, some modules may be implemented by the same physical entity, or some modules may be implemented by multiple physical entities, or they may be jointly implemented by certain components in multiple independent devices.
[0067] The above-described embodiments are merely preferred embodiments provided to fully illustrate the present invention, and the scope of protection of the present invention is not limited thereto. Equivalent substitutions or modifications made by those skilled in the art based on the present invention are all within the scope of protection of the present invention. The scope of protection of the present invention is defined by the claims.
Claims
1. A method for managing computing power in distributed AI training tasks in a mobile network, characterized by: include: Constructing an integrated mobile private network combining sensing, computing, and computing: Adding converged computing and network sensing network elements, converged computing and network control and scheduling network elements, and converged computing and network service orchestration network elements to the core network. The converged computing and network sensing network elements perceive the status of network and computing resources; the converged computing and network control and scheduling network elements control the scheduling strategies for network and computing resources; and the converged computing and network service orchestration network elements orchestrate distributed AI training tasks. When a mobile terminal accesses the integrated mobile private network for communication, sensing, and computing, a session is established for the mobile terminal's AI training task through the core network. The network elements, through the converged computing and network service orchestration, break down the AI training task into distributed collaborative training sub-tasks for computation based on the computing resources of the mobile terminal, base station, and cloud node. When a mobile terminal triggers a handover, the core network identifies the network handover signaling and executes the corresponding AI training task's computing power service handover according to the network handover signaling. The computing power service handover includes: identifying the network handover signaling Path Switch Request through the core network, executing the XN handover of the mobile terminal's computing power service according to the network handover signaling, causing a change in computing power nodes. Specifically, after the core network identifies the network handover signaling Path Switch Request, it switches the user SUPI, target base station Target-RAN-ID, and handover type Xn through the computing network convergence service orchestration network element. The computing network convergence service orchestration network element finds the distributed AI training task identifier TASKID based on the user SUPI and updates the computing power nodes and training strategy of the AI training task. When the Xn handover begins, it stops updating parameters with other computing power nodes, notifies the target base station computing power node to reserve computing power resources and preheat the model, and synchronizes the old computing power node parameters to the new computing power node through the Xn interface between the source base station and the target base station. After the handover, the new computing power node resumes computing based on the existing training experience of the old computing power node and performs collaborative training synchronization with the cloud node.
2. The computing power management method for distributed AI training tasks in a mobile network according to claim 1, characterized in that: The switching of computing power services also includes: The core network identifies the network handover signaling Handover Required and performs N2 handover of the computing power service of the mobile terminal according to the network handover signaling, causing a change in the computing power node.
3. The computing power management method for distributed AI training tasks in a mobile network according to claim 2, characterized in that: Switching to N2 causes changes in computing power nodes, including: When the core network identifies the HandOver Requested network handover signaling, it switches the user SUPI, target base station Target-RAN-ID, and handover type N2 through the computing-network convergence service orchestration network element. The computing-network convergence service orchestration network element finds the distributed AI training task identifier TASKID based on the user SUPI and updates the computing power nodes and training strategy of the AI training task. During the N2 handover preparation phase, it notifies the target base station computing power nodes to reserve computing power resources and warm up the model. During the handover, it stops updating parameters with other computing power nodes and synchronizes the old computing power node parameters to the new computing power node through the transmission tunnel between the source base station, the core network UPF, and the target base station. After the handover, the new computing power node resumes computing based on the existing training experience of the old computing power node and performs collaborative training synchronization with the cloud node.
4. A computing power management device for distributed AI training tasks in a mobile network, characterized in that: Includes a mobile private network management module and a computing power management module. The mobile private network management module constructs an integrated mobile private network combining sensing, computing, and computing: It adds converged computing and network sensing network elements, converged computing and network control and scheduling network elements, and converged computing and network service orchestration network elements to the core network side. The converged computing and network sensing network elements perceive the status of network and computing resources; the converged computing and network control and scheduling network elements control the scheduling strategies for network and computing resources; and the converged computing and network service orchestration network elements orchestrate distributed AI training tasks. When a mobile terminal accesses the integrated mobile private network, the core network establishes a session for the mobile terminal's AI training task through the computing power management module. The computing power management module, through network elements that orchestrate computing and network convergence services, breaks down the AI training task into distributed collaborative training sub-tasks for computation based on the computing resources of the mobile terminal, base station, and cloud node. When a mobile terminal triggers a handover, the computing power management module identifies the network handover signaling through the core network and executes the corresponding AI training task's computing power service handover according to the network handover signaling. The computing power management module executes the computing power service handover, including: identifying the network handover signaling Path Switch Request through the core network, and executing the XN handover of the mobile terminal's computing power service according to the network handover signaling, causing a change in computing power nodes. Specifically, after the core network identifies the network handover signaling Path Switch Request, it switches the user SUPI, target base station Target-RAN-ID, and handover type Xn through the computing network convergence service orchestration network element. It finds the distributed AI training task identifier TASKID based on the user SUPI through the computing network convergence service orchestration network element, and updates the computing power nodes and training strategy of the AI training task. When the Xn handover begins, it stops updating parameters with other computing power nodes, notifies the target base station computing power node to reserve computing power resources and preheat the model, and synchronizes the old computing power node parameters to the new computing power node through the Xn interface between the source base station and the target base station. After the handover, the new computing power node resumes computing based on the existing training experience of the old computing power node and performs collaborative training synchronization with the cloud node.
5. The computing power management device for distributed AI training tasks in a mobile network according to claim 4, characterized in that: The computing power management module performs computing power service switching, and also includes: The core network identifies the network handover signaling Handover Required and performs N2 handover of the computing power service of the mobile terminal according to the network handover signaling, causing a change in the computing power node.
6. The computing power management device for distributed AI training tasks in a mobile network according to claim 5, characterized in that: The computing power management module performs an N2 switch, causing changes in computing power nodes, including: When the core network identifies the HandOver Requested network handover signaling, it switches the user SUPI, target base station Target-RAN-ID, and handover type N2 through the computing-network convergence service orchestration network element. The computing-network convergence service orchestration network element finds the distributed AI training task identifier TASKID based on the user SUPI and updates the computing power nodes and training strategy of the AI training task. During the N2 handover preparation phase, it notifies the target base station computing power nodes to reserve computing power resources and warm up the model. During the handover, it stops updating parameters with other computing power nodes and synchronizes the old computing power node parameters to the new computing power node through the transmission tunnel between the source base station, the core network UPF, and the target base station. After the handover, the new computing power node resumes computing based on the existing training experience of the old computing power node and performs collaborative training synchronization with the cloud node.
Citation Information
Patent Citations
Equipment switching method and device, equipment and readable storage medium
CN114128349A
Method, user equipment and access network node
WO2024176812A1