Disaster recovery backup method and device, nonvolatile storage medium and electronic equipment
By configuring user plane function network elements in the disaster recovery backup system and using the session management function unit for fault detection and switching, cross-DC level disaster recovery and load sharing are achieved, solving the problems of single point of failure and low resource utilization of UPF, and improving the stability of 5G network and the utilization rate of UPF.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA TELECOM CORP LTD
- Filing Date
- 2024-12-23
- Publication Date
- 2026-05-19
AI Technical Summary
Existing UPF deployments have a single point of failure risk and are insufficient in cross-data center level disaster recovery, resulting in low resource utilization.
By configuring user plane function network elements in the disaster recovery backup system and using the session management function unit for fault detection and switching, cross-DC level disaster recovery and load sharing are achieved. The centralized shared UPF mode is used to interconnect with the campus-based UPF. The keep-alive mechanism is implemented using GRE tunnels and SMF network element point detection, and automatic one-click switching to a remote UPF is achieved.
It improved the stability of 5G networks and the utilization rate of UPF, solved the potential for single-point failures, achieved cross-DC level disaster recovery and load sharing, and enhanced network robustness and resource utilization.
Smart Images

Figure CN119814532B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, and more specifically, to a disaster recovery backup method, apparatus, non-volatile storage medium, and electronic device. Background Technology
[0002] With the rapid development of 5G networks, User Plane Functions (UPF), as a key component of the 5G core network, is responsible for packet routing and forwarding, and has extremely high requirements for service continuity. However, existing UPF deployments have the risk of single points of failure and are insufficient in terms of cross-data center (DC) level disaster recovery.
[0003] There is currently no effective solution to the above problems. Summary of the Invention
[0004] This application provides a disaster recovery backup method, apparatus, non-volatile storage medium, and electronic device to at least solve the technical problems of single-point failure risks and low resource utilization caused by the single-point deployment of user plane functions in the prior art.
[0005] According to one aspect of the embodiments of this application, a disaster recovery backup method is provided, comprising: a session management function unit in a disaster recovery backup system configuring the service range of a first user plane function server and a second user plane function server of a user plane function network element according to preset configuration information, wherein the user plane function network element is set in the disaster recovery backup system, the preset configuration information includes a tracking area code of a preset park, the user plane function network element is used as a disaster recovery backup user plane function network element for each park within the service range, and is used to process sessions within each park; after the session management function unit detects a failure of the first user plane function server, it deletes the session information of a first user session carried by the first user plane function server and instructs the first user session to transmit data through the second user plane function server, wherein the first user session is a session within the service range; after detecting that the first user plane function server has recovered from the failure, and upon receiving a second user session in which the second user session specifies the first user plane function server as a diversion anchor point, it instructs the second user session to transmit data through the first user plane function server.
[0006] Optionally, the disaster recovery backup system also includes a first service switch, a second service switch, and a management switch. The first user plane function server and the second user plane function server are connected via a heartbeat interface. The first user plane function server and the first service switch are connected via a first preset interface, and the second user plane function server and the second service switch are connected via a second preset interface. The interface types of the first preset interface include N3, N6, N4, and N9, and the types of the second preset interface also include N3, N6, N4, and N9. The first user plane function server and the management switch are connected via a third preset interface, and the second user plane function server and the management switch are connected via a fourth preset interface. The interface types of the third preset interface include an intelligent platform management interface and an operation management and maintenance interface, and the interface types of the fourth preset interface also include an intelligent platform management interface and an operation management and maintenance interface.
[0007] Optionally, the disaster recovery backup system also includes multiple bearer network devices, wherein the bearer network devices are connected to the first service switch and the second service switch via Ethernet; the first service switch, the second service switch and the management switch are all connected via Ethernet; and the first service switch and the second service switch are connected via Ethernet.
[0008] Optionally, the service area includes Class I and Class II campuses. The user plane function network elements are interconnected with the user plane function network elements of each Class I campus via dedicated lines. The user plane function network elements are used to handle sessions between the user plane function network elements of each Class I campus, and the user plane network elements of each Class I campus are subordinate nodes of the user plane function network elements. The user plane function network elements are interconnected with the user plane function network elements of each Class II campus via dedicated lines. The user plane function network elements are used as disaster recovery backup user plane function network elements for the user plane function network elements of the Class II campus.
[0009] Optionally, the service escape channel between the user plane functional network elements in the second campus is composed of the user plane functional network elements in the second campus, the first network switching device in the disaster recovery backup system, the first and second routes in the bearer network, the second network switching device in the disaster recovery backup system, and the user plane functional network elements connected in sequence.
[0010] Optionally, the capacity of the user plane function network elements in the disaster recovery backup system includes basic capacity and disaster recovery capacity. The basic capacity is equal to the sum of the capacities of all user plane function network elements in the first type of campus, and the disaster recovery capacity is equal to the sum of the capacities of all user plane function network elements in the second type of campus.
[0011] Optionally, the disaster recovery backup method further includes: determining a set of preset number of parks corresponding to the disaster recovery backup system, wherein the set of preset number of parks includes multiple alternative preset number of parks, and the preset number of parks is the number of parks sharing a single disaster recovery backup system; determining the failure conflict probability corresponding to each preset number of parks, wherein the failure conflict probability is the probability that at least two parks among multiple parks corresponding to the same disaster recovery backup system will fail simultaneously, and the failure conflict probability is used to determine the target number of parks from the set of preset number of parks.
[0012] According to another aspect of the embodiments of this application, a disaster recovery backup device is also provided, comprising: a first processing module, configured to configure the service range of a first user plane function server and a second user plane function server of a user plane function network element according to preset configuration information, wherein the preset configuration information includes a tracking area code of a preset park, the user plane function network element is used as a disaster recovery backup user plane function network element for each park within the service range, and is used to process sessions within each park; a second processing module, configured to delete the session information of a first user session carried by the first user plane function server after the session management function unit detects a failure of the first user plane function server, and instruct the first user session to transmit data through the second user plane function server, wherein the first user session is a session within the service range; and a third processing module, configured to, after detecting that the first user plane function server has recovered from a failure, and upon receiving a second user session in which the second user session specifies the first user plane function server as a diversion anchor point, instruct the second user session to transmit data through the first user plane function server.
[0013] According to another aspect of the embodiments of this application, a non-volatile storage medium is also provided, wherein a program is stored in the non-volatile storage medium, wherein the program controls the device where the non-volatile storage medium is located to perform a disaster recovery backup method when it runs.
[0014] According to another aspect of the embodiments of this application, an electronic device is also provided, including: a memory and a processor, wherein the processor is configured to run a program stored in the memory, wherein the program executes a disaster recovery backup method during runtime.
[0015] According to another aspect of the embodiments of this application, a computer program product is also provided, including a computer program that implements a disaster recovery backup method when executed by a processor.
[0016] In this embodiment, the session management function unit in the disaster recovery backup system configures the service range of the first user plane function server and the second user plane function server of the user plane function network element according to preset configuration information. The user plane function network element is located in the disaster recovery backup system, and the preset configuration information includes the tracking area code of a preset campus. The user plane function network element serves as the disaster recovery backup user plane function network element for each campus within its service range, and is used to handle sessions within each campus. After the session management function unit detects a failure in the first user plane function server, it deletes the session information of the first user session carried by the first user plane function server and instructs the first user session... Data is transmitted through the second user plane function server, where the first user session is a session within the service range; after the first user plane function server is detected to have recovered from a fault, and upon receiving the second user session and specifying the first user plane function server as the offloading anchor point in the second user session, the second user session is instructed to transmit data through the first user plane function server. By deploying a centralized shared UPF, the goal of achieving cross-DC level disaster recovery and load sharing is achieved, thereby improving the technical effect of 5G network stability and UPF utilization. This solves the technical problems of single-point failure risk and low resource utilization caused by the single-point deployment of user plane functions in the existing technology. Attached Figure Description
[0017] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments of this application and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0018] Figure 1 This is a schematic diagram of the structure of a computer terminal according to an embodiment of this application;
[0019] Figure 2 This is a flowchart illustrating a disaster recovery backup method provided according to an embodiment of this application;
[0020] Figure 3 This is a schematic diagram of the architecture of an optional user plane function network element according to an embodiment of this application;
[0021] Figure 4 This is a schematic diagram of an optional disaster recovery topology according to an embodiment of this application;
[0022] Figure 5 This is a schematic diagram of an optional GRE tunnel according to an embodiment of this application;
[0023] Figure 6 This is a schematic diagram of a disaster recovery backup device provided according to an embodiment of the present invention. Detailed Implementation
[0024] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0026] To better understand the embodiments of this application, the technical terms involved in the embodiments of this application are explained below:
[0027] 5G UPF (User Plane Function): The 5G UPF is a key user plane network element in the 5G core network, responsible for packet routing and forwarding, supporting edge computing and network slicing, and enabling low-latency, high-bandwidth services. As the anchor point for the connection between 5G and MEC (Multi-access Edge Computing), it can be flexibly deployed at the network edge, promoting local data offloading and efficient processing. Controlled by the SMF (Session Management Function), the UPF executes policies to ensure accurate data transmission, serving as the cornerstone of efficient data processing and widespread application in 5G networks, and driving the in-depth development of digital transformation.
[0028] Disaster Tolerance: Disaster tolerance refers to maintaining the uninterrupted operation of the survival system while minimizing the loss of data in the production system during accidents such as natural disasters, equipment failures, and human-caused damage.
[0029] Cross-DC disaster recovery: This refers to disaster recovery across data centers, which means establishing backup data centers in different geographical locations to ensure that data and business can be quickly restored in the event of a disaster.
[0030] As the scale of 5G customized network construction gradually expands, the UPF (User Plane Function), as a core network element for forwarding data traffic, must maintain high reliability to meet service continuity requirements. Currently, most UPF devices deployed in networks are single points of failure, posing a significant risk and presenting a large number of such devices. Adopting a 1:1 primary / backup ratio would greatly increase maintenance and investment, necessitating customized and reliable disaster recovery methods to ensure service continuity in the event of a failure. Furthermore, due to the limited disaster recovery levels of network elements, a dual-machine approach is typically used for disaster recovery: two devices are used for backup, resulting in low equipment utilization. Moreover, deploying two devices in a single data center does not meet the requirements for cross-data center (DC) and cross-network element level disaster recovery. Therefore, a more convenient and precise method for UPF disaster recovery backup is needed in practice.
[0031] To address the aforementioned issues, this application provides relevant solutions, which are detailed below.
[0032] According to an embodiment of this application, a method embodiment of a disaster recovery backup method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0033] The methods and embodiments provided in this application can be executed on mobile terminals, computer terminals, or similar computing devices. Figure 1 A hardware block diagram of a computer terminal (or mobile device) for implementing a disaster recovery backup method is shown. Figure 1 As shown, the computer terminal 10 (or mobile device 10) may include one or more processors 102 (shown as 102a, 102b, ..., 102n in the figure) 102 (processor 102 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 10 may also include... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0034] It should be noted that the aforementioned one or more processors 102 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 10 (or mobile device). As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0035] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the disaster recovery backup method in this embodiment. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, thereby realizing the disaster recovery backup method described above. The memory 104 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the computer terminal 10 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0036] The transmission device 106 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 10. In one example, the transmission device 106 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 106 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0037] The display can be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 10 (or mobile device).
[0038] Under the above operating environment, this application provides a disaster recovery backup method, such as... Figure 2 As shown, the method includes the following steps:
[0039] In step S202, the session management function unit in the disaster recovery backup system configures the service range of the first user plane function server and the second user plane function server of the user plane function network element according to the preset configuration information. The user plane function network element is set in the disaster recovery backup system. The preset configuration information includes the tracking area code of the preset campus. The user plane function network element is used as the disaster recovery backup user plane function network element of each campus within the service range, and is used to handle sessions within each campus.
[0040] Optionally, the user plane function network elements are deployed using a load-sharing approach, with UPF-1 (the first user plane function server) and UPF-2 (the second user plane function server) forming a service pool. Both UPF servers are connected to the enterprise intranet. Specific service configurations for the SMF (Session Management Function) and the switch include:
[0041] 1) SMF: Configure the service range of both UPF-1 and UPF-2 to be the tracking area code of the campus. UPF-1 and UPF-2 are configured with the same UE (User Equipment) external address pool and are configured to publish UE downlink routes using dynamic routing protocol.
[0042] 2) Switch: Configure normal network topology between the switch and UPF1 / 2, and configure dynamic routing protocol;
[0043] Optionally, if the switch connects to the enterprise intranet using the GRE protocol and the GRE protocol endpoint is on the UPF, then only the primary / standby mode can be used.
[0044] In the technical solution provided in step S202, the disaster recovery backup system also includes a first service switch, a second service switch, and a management switch. The first user plane function server and the second user plane function server are connected via a heartbeat interface. The first user plane function server and the first service switch are connected via a first preset interface, and the second user plane function server and the second service switch are connected via a second preset interface. The interface types of the first preset interface include N3, N6, N4, and N9, and the types of the second preset interface include N3, N6, N4, and N9. The first user plane function server and the management switch are connected via a third preset interface, and the second user plane function server and the management switch are connected via a fourth preset interface. The interface types of the third preset interface include an intelligent platform management interface and an operation management and maintenance interface, and the interface types of the fourth preset interface include an intelligent platform management interface and an operation management and maintenance interface.
[0045] In the technical solution provided in step S202, the disaster recovery backup system also includes multiple bearer network devices, wherein the bearer network devices are connected to the first service switch and the second service switch via Ethernet; the first service switch, the second service switch and the management switch are all connected via Ethernet; and the first service switch and the second service switch are connected via Ethernet.
[0046] Optionally, a centralized shared UPF (i.e., user plane function network element in the disaster recovery backup system) is deployed in the city as the business anchor point for Type A campus (a dedicated campus in an enterprise) (i.e., the user plane function network element is used to handle sessions within each Type A campus), and also serves as the disaster recovery UPF for the campus (i.e., as the disaster recovery backup user plane function network element for a campus located in a different location from the shared UPF in another location). Figure 3 The architecture diagram of the user plane function network elements is shown, such as... Figure 3 As shown, the UPF network element, as a user plane control network element, consists of 2 * UPF servers, 2 * service switches, and 1 * management switch. The UPF adopts an m-lag or stacking method. Server 1 (first user plane function server) - Server 2 (second user plane function server): Data synchronization and heartbeat detection are achieved through a heartbeat interface to ensure high availability of the two UPF servers. Service switch - server: interconnected through N3 / N6 / N4 / N9 VLAN configuration (i.e., the first user plane function server and the first service switch are connected through the first preset interface, and the second user plane function server and the second service switch are connected through the second preset interface). The N3 / N6 ports can achieve traffic load sharing through interface bonding. The N6 port can implement multi-customer multi-DN (data network) mode through physical interfaces or sub-interfaces. Server - managed switch; OAM (Operation, Administration, and Maintenance) and IPMI (Intelligent Management Interface). The Platform Management Interface (PMI) is used for access (i.e., the first user plane function server and the management switch are connected via the third preset interface, and the second user plane function server and the management switch are connected via the fourth preset interface). Management switch - service switch: Outband VLAN interconnection. Refer to the table below for interface-related descriptions. All interfaces are connected via Ethernet.
[0047]
[0048]
[0049] Step S204: After the session management function unit detects a failure in the first user plane function server, it deletes the session information of the first user session carried by the first user plane function server and instructs the first user session to transmit data through the second user plane function server, wherein the first user session is a session within the service range.
[0050] Optionally, under normal circumstances, SMF will round-robin the user sessions to UPF-1 (first user plane function server) and UPF-2 (second user plane function server), and the user will access the enterprise intranet through the corresponding UPF.
[0051] After the SMF detects a fault in the splitter UPF-1, it deletes the user session (the first user session) on UPF-1 and then inserts UPF-2 as an ULCL (Uplink Classifier) into the user session. Users then access the enterprise intranet through UPF-2. The user remains online during the handover process, and there is no signaling impact on surrounding network elements (such as PCFs).
[0052] Step S206: After detecting that the first user plane function server has recovered from a failure, and upon receiving a second user session in which the second user session specifies the first user plane function server as the routing anchor, instruct the second user session to transmit data through the first user plane function server.
[0053] Optionally, when the UPF-1 fault is recovered, the new user session (the second user session) can choose this UPF-1 as the offloading anchor point. The second user session transmits data through the first user plane function server, and online users will not be migrated to UPF1.
[0054] Optionally, the service area includes Class I and Class II campuses. The user plane function network elements are interconnected with the user plane function network elements of each Class I campus via dedicated lines. The user plane function network elements are used to handle sessions between the user plane function network elements of each Class I campus, and the user plane network elements of each Class I campus are subordinate nodes of the user plane function network elements. The user plane function network elements are interconnected with the user plane function network elements of each Class II campus via dedicated lines. The user plane function network elements are used as disaster recovery backup user plane function network elements for the user plane function network elements of the Class II campus.
[0055] Optionally, Figure 4 A disaster recovery topology diagram is shown, such as Figure 4As shown, the UPFx (user plane network element) of Park A (Category 1 Park) acts as a subordinate node of the user plane function network element (UPF1). The city-level shared UPF1 (user plane function network element) is interconnected with all UPFs (user plane network elements) of Park A via dedicated lines for shared service use. The shared UPF1 (user plane function network element) acts as the primary control anchor UPF and is interconnected with the UPF (user plane function network element) of Park Type B (Category 2 Park) with disaster recovery relationships via dedicated lines as its backup (disaster recovery backup user plane function network element). The core network SMF network element point is used to detect the network connectivity of the primary and backup nodes to achieve a keep-alive mechanism. Among them, the user plane function network element is located in a different location from the Category 2 Park, thereby achieving off-site disaster recovery. Park Type A: City-level shared UPF is selected, that is, the business traffic within the park is processed through the city-level UPF. Park Type B: In-situ UPF is deployed, that is, the UPF is deployed directly within the park to reduce the transmission distance of business traffic and improve efficiency. The green dashed line represents the normal business flow of the UPF located in the park, that is, the business traffic is processed within the park. The red dashed line represents the business flow when the UPF located in the park fails, that is, when the UPF in the park fails, the business traffic will be processed through the shared UPF1 (user plane function network element) at the city level.
[0056] Optionally, the service escape channel between the user plane functional network elements in the second campus is composed of the user plane functional network elements in the second campus, the first network switching device in the disaster recovery backup system, the first and second routes in the bearer network, the second network switching device in the disaster recovery backup system, and the user plane functional network elements connected in sequence.
[0057] Optionally, the core + bearer joint disaster recovery specifically includes: the main UPF (user plane functional network element of the second campus) combined with the uplink bearer network B device, using VPN based on GRE to establish a tunnel from the main UPF to network switch 1 (first network switch) to B1 (first route) to B2 (second route) to network switch 2 (second network switch) to UPF1 (user plane functional network element) as a service escape channel. The SMF sets the service routing priority cost value to prioritize service access from the higher-priority UPF. The service escape channel is a network design strategy used to ensure service continuity and uninterrupted data transmission when the main network path or equipment fails. Through VPN (Virtual Private Network) and GRE (Generic Routing Encapsulation) technologies, a backup communication tunnel can be established, allowing service traffic to switch to this escape channel when the main communication link fails.
[0058] The detection mechanism and recovery principle specifically include: the core network SMF element detects the UPF keep-alive mechanism through the N4 interface. When the main UPF is detected as unreachable, traffic is automatically switched to the shared UPF (UPF1) through the GRE tunnel. After the faulty main UPF is restored, traffic is automatically switched to the main UPF, and services are seamlessly migrated online without being affected.
[0059] Optionally, the capacity of the user plane function network element in the disaster recovery backup system includes basic capacity and disaster recovery capacity. The basic capacity is equal to the sum of the capacities of all user plane function network elements in the first type of campus, and the disaster recovery capacity is equal to the maximum value among the capacities of all user plane function network elements in the second type of campus.
[0060] Optionally, each campus uses a dedicated DNN (Data Network Name). Under the above deployment method, the capacity of the user plane function network elements is determined as follows:
[0061] 1) The capacity of the centralized shared UPF (User Plane Functional Network Element) in the city is divided into the basic capacity and the disaster recovery capacity. The basic capacity is the sum of the capacities of the user plane functional network elements in each of the first-class parks. The disaster recovery capacity is the sum of the capacities of the user plane functional network elements in each of the second-class parks. If there are multiple first-class parks, the basic capacity is the sum of the capacities of the user plane functional network elements in each of the first-class parks. If there are multiple second-class parks, the disaster recovery capacity is the maximum value of the sum of the capacities of the user plane functional network elements in each of the second-class parks.
[0062] 2) The basic capacity of the shared UPF in the city cannot be 0. There must be business traffic on a regular basis to ensure that the equipment and network environment are normal.
[0063] 3) Basic capacity and disaster recovery capacity are separated and isolated, and uniformly planned on the city-shared UPF (User Plane Functional Element); when the capacity of the city-shared UPF is insufficient, it needs to be expanded in pairs. Specifically, the city-shared UPF (User Plane Functional Element) is set up in pairs, such as... Figure 4 The system consists of two UPF1 and UPF2 network elements. During expansion, the two cities share the same UPF and expand synchronously.
[0064] Optionally, the disaster recovery backup method further includes: determining a set of preset number of parks corresponding to the disaster recovery backup system, wherein the set of preset number of parks includes multiple alternative preset number of parks, and the preset number of parks is the number of parks sharing a single disaster recovery backup system; determining the failure conflict probability corresponding to each preset number of parks, wherein the failure conflict probability is the probability that at least two parks among multiple parks corresponding to the same disaster recovery backup system will fail simultaneously, and the failure conflict probability is used to determine the target number of parks from the set of preset number of parks.
[0065] Optionally, a convergence ratio analysis is performed on the number of disaster recovery sets for centralized shared UPFs, specifically including:
[0066] 1) The availability of a centralized shared UPF is 99.9%, which means an average downtime of 530 minutes per year;
[0067] 2) Based on the impact of UPF failures, the average failure duration needs to be mapped to the average number of failures per year.
[0068] 3) If the UPF fails an average of 1 time per year, then the probability of a failure conflict for n sets of UPFs (n≤365):
[0069]
[0070] 4) If the average number of UPF failures per year is 2, then the probability of a failure conflict for n sets of UPFs (n≤182):
[0071]
[0072] 5) If the average number of UPF failures is 3 per year, then the probability of a failure conflict for n sets of UPFs (n≤121) is:
[0073]
[0074] The following table shows the probability of fault conflict for multiple candidate preset park numbers in the preset park number set:
[0075]
[0076] Based on the convergence ratio analysis, a suggested value of 8 for the convergence ratio N is obtained; the method for setting the suggested value of the convergence ratio N is analyzed as follows:
[0077] 1) As can be seen from the table above, if the UPF fails an average of 1 time per year, the probability of conflict is more than 50% when 23 disaster recovery sets are installed in the UPF; if the average failure is 2 times per year, the probability of conflict is more than 50% when 17 disaster recovery sets are installed; and if the average failure is 3 times per year, the probability of conflict is more than 50% when 14 disaster recovery sets are installed.
[0078] 2) Taking all factors into consideration, it is recommended that the convergence ratio N=8, that is, one disaster recovery resource is reserved on the shared UPF (User Plane Functional Element) for every 8 campuses, so as to achieve a balance between resource utilization and reliability.
[0079] Shared UPF2 can also be deployed in the same way as shared UPF1 as a centralized shared UPF node in other parks, and the implementation method is the same as that of shared UPF1.
[0080] Under the above deployment method, analysis is performed on the business disaster recovery data stream, such as... Figure 4 As shown:
[0081] The shared UPF at the city level serves as the dedicated line business flow anchor for Type A parks, and also functions as a disaster recovery UPF for Type B parks.
[0082] 1. Business traffic in Park 1 and Park 2 is shared via UPF1 and UPF2. Parks 1 and 2 are non-enterprise parks that directly use shared UPFs for business processing. Type A Parks: City-level shared UPFs are selected, meaning business traffic within the park is processed through a city-level UPF. Type B Parks: In-place UPFs are deployed, meaning UPFs are deployed directly within the park to reduce the transmission distance of business traffic and improve efficiency.
[0083] 1) When shared UPF1 fails, services in Park 1 and Park 2 will use shared UPF2;
[0084] 2) When shared UPF2 fails, services in Park 1 and Park 2 will use shared UPF1;
[0085] 2. Share UPF1 as the disaster recovery UPF for Campus A UPFx; UPF1 and Campus A UPFx are connected via a GRE tunnel:
[0086] 1) Normal traffic in Park A: forwarded through Park A UPFx;
[0087] 2) When UPFx fails in Campus A, Campus A forwards data via shared UPF1;
[0088] 3) When UPFx fails in campus A, traffic in campus A is forwarded through shared UPF1;
[0089] 4) Once the UPFx fault in the park is resolved, the traffic of existing online users will be forwarded from UPF1 to UPFx via the GRE tunnel.
[0090] 3. Shared UPF2 serves as the disaster recovery UPF for campus B UPFy; UPF2 and campus B UPFy are connected via a GRE tunnel:
[0091] 1) When traffic in Park B is normal: it is forwarded through Park B UPFy;
[0092] 2) When UPFy in Campus B fails, Campus B forwards the data via shared UPF2;
[0093] 3) When UPFy in Campus B fails, traffic in Campus B is forwarded through the shared UPF2;
[0094] 4) After the UPFy fault in the park is resolved, the traffic of existing online users will be forwarded from UPF2 to UPFy through the GRE tunnel.
[0095] Optionally, Figure 5 A schematic diagram of a GRE tunnel is shown, such as... Figure 5 As shown, the User Equipment (UE) connects to the network through the Radio Access Network (CDMA-RAN), which is responsible for transmitting and receiving radio signals. The UE communicates with the Access Point Node (APN) through the N3 interface. The APN is the network's access point node, responsible for handling traffic from the UE.
[0096] The APN connects to User Plane Functions (UPF) 01 and 02 via the CDMA-N6.Local interface. The CDMA-EPC (CDMA Evolution Packet Core), as part of the core network, connects to the UPF via B / CE / PE (Base Station Controller / Core Network Element / Packet Data Network Gateway) to route and forward data. The UPF is a key component in 5G networks, responsible for packet forwarding and processing.
[0097] The diagram also shows two tunnel interfaces, tunnel1 and loopback51, which are used to establish GRE tunnels between UPFs. GRE tunnels allow data packets to be transmitted between different network segments while maintaining packet integrity and security. Between UPF01 and UPF02, efficient data transmission is achieved through GRE tunnels, which transmit data via an internal VPN (Virtual Private Network) CDMA-N6.Local and a shared VPN (Virtual Private Network) CDMA-EPC (CDMA Evolution Packet Core).
[0098] Through the above steps, the centralized shared UPF mode enables uninterrupted transmission of cross-network element and DC-level business sessions, thereby improving network robustness and UPF utilization. A centralized shared UPF is deployed at the city level as a business anchor point and interconnected with the disaster recovery UPF leased line deployed in the park. A keep-alive mechanism is implemented based on GRE tunnels and SMF network element detection, and automated one-click switching to a remote UPF allows for rapid migration of services to the backup machine. The centralized shared UPF can handle basic business and also establish shared disaster recovery with the park's UPF, resolving the risks associated with single-point UPF deployment, improving UPF utilization, and achieving cost reduction and efficiency improvement. Compared to traditional single data centers, utilizing the high availability of two UPFs, it achieves cross-DC-level disaster recovery and hierarchical disaster recovery from the city level to the park level, while providing a method for calculating the number of shared disaster recovery nodes (N) that can handle the disaster recovery. The method embodiments of this application can be extended to the following scenarios: provincial-municipal 5G UPF cross-level disaster recovery; applying shared UPFs to application scenarios with multiple sub-nodes, such as industrial ports and educational campuses; scenarios where on-premises UPFs are used as master nodes and shared with other branch UPF nodes for disaster recovery; and using UPFs with low utilization (recommended computing resource utilization rate is less than 10%) as shared disaster recovery nodes. For 5G services, considering cost savings for customers, a single device is provided for disaster recovery with a shared UPF at the municipal level, ensuring business continuity and achieving high reliability. The method embodiments of this application promote the development of a product concept that transforms a single blade UPF server into a set of two UPFs for disaster recovery, proposing a new cross-level disaster recovery solution to meet the needs of high-reliability disaster recovery, which has significance for nationwide promotion. The method embodiments of this application extend to cross-DC cross-level solutions, realizing a master-slave UPF disaster recovery deployment mode. The method embodiments provided by this application have the following advantages:
[0099] 1. Centralized Shared UPF Deployment Module: Deploy a centralized shared UPF at the prefecture-level city level as a business anchor point, and at the same time assume the role of disaster recovery UPF for park-level deployment.
[0100] 2. Load Sharing Module: Implements load sharing for the UPF, forming a service pool by splitting UPF1 / 2 and connecting it to the enterprise intranet. The disaster recovery process is as follows: when the SMF detects a UPF failure, it automatically switches the user session to the backup UPF to ensure service continuity.
[0101] 3. Cross-Data Center Disaster Recovery Module: For a single data center's UPF, this module enables cross-data center level disaster recovery, improving network robustness and UPF utilization. It interconnects with shared UPFs and dedicated UPFs in the same campus network with disaster recovery relationships via dedicated lines, establishes tunnels based on GRE, and utilizes SMF network element detection to implement a keep-alive mechanism.
[0102] 4. Multi-level disaster recovery: From network element-level disaster recovery to network-level disaster recovery, on the basis of network element disaster recovery, the network can be further realized in different locations, which provides a higher level of guarantee for the stability of the 5G customized network.
[0103] 5. One-click switching module: By detecting the UPF keep-alive mechanism through the SMF network element, traffic is automatically switched to the other side through the GRE tunnel when the primary UPF is unavailable. It automatically blocks the faulty UPF network with one click, synchronously switches to a remote UPF, and quickly migrates services to the backup machine, maximizing service continuity.
[0104] 6. Dedicated DNN Module: Each campus uses a dedicated DNN to differentiate between different business traffic. The capacity of the centralized shared UPF in each city is differentiated and isolated into basic capacity and disaster recovery capacity, and planned uniformly.
[0105] 7. Convergence Ratio Analysis Module: Analyzes the availability and failure impact of centralized shared UPFs to determine the optimal balance between resource utilization and reliability.
[0106] This application provides a disaster recovery backup device. Figure 6 This is a schematic diagram of the device, as shown below. Figure 6 As shown, the device includes: a first processing module 60, configured to configure the service range of the first user plane function server and the second user plane function server of the user plane function network element according to preset configuration information, wherein the preset configuration information includes the tracking area code of the preset campus, the user plane function network element is used as the disaster recovery backup user plane function network element of each campus within the service range, and is used to process sessions within each campus; a second processing module 62, configured to delete the session information of the first user session carried by the first user plane function server after the session management function unit detects a failure of the first user plane function server, and instruct the first user session to transmit data through the second user plane function server, wherein the first user session is a session within the service range; and a third processing module 64, configured to, after detecting that the first user plane function server has recovered from the failure, and upon receiving a second user session in which the second user session specifies the first user plane function server as the diversion anchor point, instruct the second user session to transmit data through the first user plane function server.
[0107] In some embodiments of this application, the disaster recovery backup system further includes a first service switch, a second service switch, and a management switch. The first user plane function server and the second user plane function server are connected via a heartbeat interface. The first user plane function server and the first service switch are connected via a first preset interface, and the second user plane function server and the second service switch are connected via a second preset interface. The interface types of the first preset interface include N3, N6, N4, and N9, and the types of the second preset interface include N3, N6, N4, and N9. The first user plane function server and the management switch are connected via a third preset interface, and the second user plane function server and the management switch are connected via a fourth preset interface. The interface types of the third preset interface include an intelligent platform management interface and an operation management and maintenance interface, and the interface types of the fourth preset interface include an intelligent platform management interface and an operation management and maintenance interface.
[0108] In some embodiments of this application, the disaster recovery backup system further includes multiple bearer network devices, wherein the bearer network devices are connected to the first service switch and the second service switch via Ethernet; the first service switch, the second service switch, and the management switch are all connected via Ethernet; and the first service switch and the second service switch are connected via Ethernet.
[0109] In some embodiments of this application, the various parks within the service range include first-class parks and second-class parks. The user plane function network elements are interconnected with the user plane function network elements of each first-class park via dedicated lines. The user plane function network elements are used to process sessions between the user plane function network elements of each first-class park, and the user plane network elements of the first-class parks are subordinate nodes of the user plane function network elements. The user plane function network elements are also interconnected with the user plane function network elements of each second-class park via dedicated lines. The user plane function network elements serve as disaster recovery backup user plane function network elements for the user plane function network elements of the second-class parks.
[0110] In some embodiments of this application, the service escape channel between the second campus user plane function network element and the user plane function network element is composed of the second campus user plane function network element, the first network switching device in the disaster recovery backup system, the first route and the second route in the bearer network, and the second network switching device in the disaster recovery backup system and the user plane function network element connected in sequence.
[0111] In some embodiments of this application, the capacity of the user plane function network element in the disaster recovery backup system includes a basic capacity and a disaster recovery capacity, wherein the basic capacity is equal to the sum of the capacities of each user plane function network element in the first type of campus, and the disaster recovery capacity is equal to the sum of the capacities of each user plane function network element in the second type of campus.
[0112] In some embodiments of this application, the disaster recovery backup method further includes: determining a preset set of the number of parks corresponding to the disaster recovery backup system, wherein the preset set of the number of parks includes multiple alternative preset park numbers, and the preset park number is the number of parks sharing a single disaster recovery backup system; determining the fault conflict probability corresponding to each preset park number, wherein the fault conflict probability is the probability that at least two parks among multiple parks corresponding to the same disaster recovery backup system will fail simultaneously, and the fault conflict probability is used to determine the target park number from the preset park number set.
[0113] It should be noted that each module in the disaster recovery backup device can be a program module (e.g., a set of program instructions to implement a specific function) or a hardware module. For the latter, it can take the following forms, but is not limited to them: each of the above modules is represented by a processor, or the functions of each of the above modules are implemented by a processor.
[0114] This application provides a non-volatile storage medium storing a program. During program execution, the device containing the non-volatile storage medium performs the following disaster recovery backup method: A session management function unit in the disaster recovery backup system configures the service range of a first user plane function server and a second user plane function server within the user plane function network element, based on preset configuration information. The user plane function network element is located in the disaster recovery backup system. The preset configuration information includes a tracking area code for a preset campus. The user plane function network element serves as the disaster recovery backup user plane function network element for each campus within its service range and is used to handle sessions within each campus. After the session management function unit detects a failure in the first user plane function server, it deletes the session information of the first user session carried by the first user plane function server and instructs the first user session to transmit data through the second user plane function server. The first user session is a session within the service range. After detecting that the first user plane function server has recovered from the failure, and upon receiving a second user session that specifies the first user plane function server as a diversion anchor point, the second user session is instructed to transmit data through the first user plane function server.
[0115] This application provides an electronic device, including a memory and a processor. The processor is used to run a program stored in the memory. When the program runs, it executes the following disaster recovery backup method: A session management function unit in the disaster recovery backup system configures the service range of a first user plane function server and a second user plane function server of a user plane function network element according to preset configuration information. The user plane function network element is set in the disaster recovery backup system. The preset configuration information includes a tracking area code for a preset campus. The user plane function network element is used as a disaster recovery backup user plane function network element for each campus within its service range, and is used to process sessions within each campus. After the session management function unit detects a failure in the first user plane function server, it deletes the session information of the first user session carried by the first user plane function server and instructs the first user session to transmit data through the second user plane function server. The first user session is a session within the service range. After detecting that the first user plane function server has recovered from the failure, and upon receiving a second user session in which the second user session specifies the first user plane function server as a diversion anchor point, it instructs the second user session to transmit data through the first user plane function server.
[0116] This application provides a computer program product, including a computer program that, when executed by a processor, implements the following disaster recovery backup method: A session management function unit in the disaster recovery backup system configures the service range of a first user plane function server and a second user plane function server of a user plane function network element according to preset configuration information. The user plane function network element is located in the disaster recovery backup system. The preset configuration information includes a tracking area code for a preset campus. The user plane function network element is used as a disaster recovery backup user plane function network element for each campus within its service range, and is used to handle sessions within each campus. After the session management function unit detects a failure in the first user plane function server, it deletes the session information of the first user session carried by the first user plane function server and instructs the first user session to transmit data through the second user plane function server. The first user session is a session within the service range. After detecting that the first user plane function server has recovered from the failure, and upon receiving a second user session in which the second user session specifies the first user plane function server as a diversion anchor point, it instructs the second user session to transmit data through the first user plane function server.
[0117] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0118] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0119] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0120] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0121] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to related technologies, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0122] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A disaster recovery backup method, characterized in that, include: The session management function unit in the disaster recovery backup system configures the service range of the first user plane function server and the second user plane function server of the user plane function network element according to preset configuration information. The user plane function network element is located in the disaster recovery backup system. The preset configuration information includes the tracking area code of a preset park. The user plane function network element serves as the disaster recovery backup user plane function network element for each park within the service range, and is used to handle sessions within each park. The parks within the service range include first-class parks and second-class parks. The user plane function network element is interconnected with the user plane function network elements of each first-class park in the first-class parks via dedicated lines. The user plane function element is used to handle user plane data in each first-class park. The user plane functional network element is interconnected with each user plane functional network element in the second type of campus via a dedicated line. The user plane functional network element is used as a disaster recovery backup user plane functional network element for the user plane functional network elements in the second type of campus. The capacity of the user plane functional network element in the disaster recovery backup system includes a basic capacity and a disaster recovery capacity. The basic capacity is equal to the sum of the capacities of each user plane functional network element in the first type of campus, and the disaster recovery capacity is equal to the maximum value of the capacities of each user plane functional network element in the second type of campus. If the capacity of the user plane functional network element is insufficient, the first user plane functional server and the second user plane functional server in the user plane functional network element are expanded synchronously. After the session management function unit detects a failure in the first user plane function server, it deletes the session information of the first user session carried by the first user plane function server and instructs the first user session to transmit data through the second user plane function server, wherein the first user session is a session within the service scope. After detecting that the first user plane function server has recovered from a failure, and upon receiving a second user session in which the second user session specifies the first user plane function server as the routing anchor, the second user session is instructed to transmit data through the first user plane function server.
2. The disaster recovery backup method according to claim 1, characterized in that, The disaster recovery backup system also includes a first service switch, a second service switch, and a management switch, wherein... The first user plane function server and the second user plane function server are connected via a heartbeat interface; The first user plane function server and the first service switch are connected through a first preset interface, and the second user plane function server and the second service switch are connected through a second preset interface. The interface types of the first preset interface include: N3 interface, N6 interface, N4 interface and N9 interface, and the types of the second preset interface include: N3 interface, N6 interface, N4 interface and N9 interface. The first user plane function server and the management switch are connected through a third preset interface, and the second user plane function server and the management switch are connected through a fourth preset interface. The interface types of the third preset interface include intelligent platform management interface and operation management and maintenance interface, and the interface types of the fourth preset interface include intelligent platform management interface and operation management and maintenance interface.
3. The disaster recovery backup method according to claim 2, characterized in that, The disaster recovery backup system also includes multiple bearer network devices, among which... The bearer network device is connected to the first service switch and the second service switch via Ethernet. The first service switch, the second service switch, and the management switch are all connected via Ethernet; The first service switch and the second service switch are connected via Ethernet.
4. The disaster recovery backup method according to claim 1, characterized in that, The service escape channel between the second campus user plane functional network element and the user plane functional network element is composed of the second campus user plane functional network element, the first network switching device in the disaster recovery backup system, the first and second routes in the bearer network, the second network switching device in the disaster recovery backup system, and the user plane functional network element connected in sequence.
5. The disaster recovery backup method according to claim 1, characterized in that, The disaster recovery backup method also includes: Determine the set of preset number of parks corresponding to the disaster recovery backup system, wherein the set of preset number of parks includes multiple alternative preset number of parks, and the preset number of parks is the number of parks sharing a single disaster recovery backup system; Determine the fault conflict probability corresponding to each of the preset number of parks, wherein the fault conflict probability is the probability that at least two parks in multiple parks corresponding to the same disaster recovery backup system will fail at the same time, and the fault conflict probability is used to determine the target number of parks from the preset number of parks.
6. A disaster recovery backup device, applicable to the session management function unit of a disaster recovery backup system, characterized in that, include: The first processing module is used to configure the service range of the first user plane function server and the second user plane function server of the user plane function network element according to the preset configuration information. The preset configuration information includes the tracking area code of the preset park. The user plane function network element is used as the disaster recovery backup user plane function network element of each park within the service range, and is used to process the sessions in each park. The second processing module is configured to, after the session management function unit detects a fault in the first user plane function server, delete the session information of the first user session carried by the first user plane function server, and instruct the first user session to transmit data through the second user plane function server. The first user session is a session within the service range, and the various campuses within the service range include first-class campuses and second-class campuses. The user plane function network element is interconnected with the user plane function network elements of each first-class campus in the first-class campuses via dedicated lines. The user plane function network element is used to process the sessions of the user plane function network elements of each first-class campus. The user plane function network element is interconnected with the first... The user plane function network elements of each second-level campus in the second-level campus are interconnected through dedicated lines. The user plane function network elements are used as disaster recovery backup user plane function network elements of the second-level campus user plane function network elements. The capacity of the user plane function network elements in the disaster recovery backup system includes basic capacity and disaster recovery capacity. The basic capacity is equal to the sum of the capacities of each first-level campus user plane function network element in the first-level campus, and the disaster recovery capacity is equal to the maximum value of the capacities of each second-level campus user plane function network element in the second-level campus. When the capacity of the user plane function network element is insufficient, the first user plane function server and the second user plane function server in the user plane function network element are expanded simultaneously. The third processing module is used to, after detecting that the first user plane function server has recovered from a failure, and upon receiving a second user session in which the second user session specifies the first user plane function server as a routing anchor, instruct the second user session to transmit data through the first user plane function server.
7. A non-volatile storage medium, characterized in that, The non-volatile storage medium stores a program, wherein when the program is executed, it controls the device where the non-volatile storage medium is located to execute the disaster recovery backup method according to any one of claims 1 to 5.
8. An electronic device, characterized in that, include: A memory and a processor, the processor being configured to run a program stored in the memory, wherein the program, when running, executes the disaster recovery backup method according to any one of claims 1 to 5.
9. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the disaster recovery backup method according to any one of claims 1 to 5.