Replacement of unstable application
Patent Information
- Application Number
- JP2024566968
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-06-25
- Estimated Expiration
- 2042-12-26
AI Technical Summary
When replacing unstable applications, existing methods often result in wastage of hardware resources as all processes are replaced, even if only some hardware resources are causing instability, leading to inefficient resource utilization.
A replacement system that determines application instability based on stability evaluation values across multiple resources, allowing for selective exclusion and addition of resources within an active resource group, and iteratively checks for stability post-configuration changes to pinpoint and isolate the cause of instability.
This approach minimizes resource wastage by accurately identifying and isolating the unstable resource, ensuring only necessary resources are replaced, thus optimizing hardware utilization.
Abstract
Description
Replacing unstable applications
[0001] The present invention relates to replacing unstable applications.
[0002] Patent Document 1 describes deploying a network function included in a communication system to a server on which a container-type application execution environment is installed.
[0003] International Publication No. 2021 / 171210
[0004] Some applications, such as network functions, that operate in a communication system include many processes, and these processes may be distributed and run on multiple hardware resources.
[0005] When evaluating the stability of such an application, determining whether each of the many processes included in the application is unstable based on a stability evaluation value indicating the stability of the process would require an enormous amount of calculation.
[0006] Therefore, when evaluating the stability of such an application, it is common to determine whether the application is unstable based on a stability evaluation value aggregated for multiple processes. For example, for each type of process included in the application, it is determined whether the application is unstable based on a stability evaluation value aggregated for multiple processes of that type.
[0007] However, if an application is determined to be unstable based on a stability evaluation value aggregated across multiple processes, since the stability evaluation value is an aggregate value across multiple hardware resources, it is possible that the cause of the application's instability lies in one of the hardware resources across which the processes included in the application are distributed and running.
[0008] In such a situation, if all processes included in the application are replaced with other hardware resources, processes running on hardware resources that are not the cause of the application's instability will also require replacement hardware resources, resulting in a waste of hardware resources.
[0009] The above is not limited to applications included in communication systems, but also applies to general applications.
[0010] The present invention has been made in view of the above-mentioned circumstances, and one of its objects is to make it possible to reduce waste of hardware resources at the replacement destination when replacing an unstable application.
[0011] In order to solve the above problem, the replacement system of the present disclosure includes an instability determination means that determines whether an application is unstable based on a stability evaluation value that indicates the stability of an application whose processes are distributed across a group of operating resources including a plurality of replacement candidate resources, and a configuration change means that, if the application is determined to be unstable based on the stability evaluation value that indicates the stability of the application whose processes are distributed across the group of operating resources, executes a configuration change that excludes at least one of the replacement candidate resources from the group of operating resources and adds a new resource to the group of operating resources; and if the application is determined to be unstable based on the stability evaluation value that indicates the stability of the application whose processes are distributed across the group of operating resources after the configuration change has been executed, executes the configuration change again.
[0012] In addition, the replacement method of the present disclosure includes determining whether an application is unstable based on a stability evaluation value indicating the stability of an application whose processes are distributed across a group of operating resources including a plurality of replacement candidate resources; if it is determined that the application is unstable based on the stability evaluation value indicating the stability of the application whose processes are distributed across the group of operating resources, executing a configuration change to exclude at least one of the replacement candidate resources from the group of operating resources and add a new resource to the group of operating resources; and if it is determined that the application is unstable based on the stability evaluation value indicating the stability of the application whose processes are distributed across the group of operating resources after the configuration change has been executed, executing the configuration change again.
[0013] 1 is a diagram illustrating an example of a communication system according to an embodiment of the present invention. FIG. 1 is a diagram illustrating an example of a communication system according to an embodiment of the present invention. FIG. 2 is a diagram illustrating an example of a network service according to an embodiment of the present invention. FIG. 3 is a diagram illustrating an example of an association between elements established in a communication system according to an embodiment of the present invention. FIG. 4 is a functional block diagram illustrating an example of functions implemented in a platform system according to an embodiment of the present invention. FIG. 5 is a diagram illustrating an example of a data structure of physical inventory data. FIG. 6 is a diagram illustrating an example of a situation in which processes included in each of a plurality of applications are distributed to run on a plurality of hardware resources. FIG. 7 is a diagram illustrating an example of a situation in which processes included in each of a plurality of applications are distributed to run on a plurality of hardware resources. FIG. 8 is a diagram illustrating an example of a situation in which processes included in each of a plurality of applications are distributed to run on a plurality of hardware resources. FIG. 9 is a diagram illustrating an example of a situation in which processes included in each of a plurality of applications are distributed to run on a plurality of hardware resources. FIG. 10 is a flow diagram illustrating an example of a flow of processing performed in a platform system according to an embodiment of the present invention.
[0014] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings.
[0015] 1 and 2 are diagrams illustrating an example of a communication system 1 according to an embodiment of the present invention. Fig. 1 is a diagram focusing on the locations of a group of data centers included in the communication system 1. Fig. 2 is a diagram focusing on various computer systems implemented in the group of data centers included in the communication system 1.
[0016] As shown in FIG. 1 , the data centers included in the communication system 1 are classified into a central data center 10 , regional data centers 12 , and edge data centers 14 .
[0017] For example, several central data centers 10 are distributed and located within the area covered by the communication system 1 (for example, within Japan).
[0018] For example, several tens of regional data centers 12 are distributed and placed within the area covered by the communication system 1. For example, if the area covered by the communication system 1 is the entire country of Japan, one or two regional data centers 12 may be placed in each prefecture.
[0019] For example, several thousand edge data centers 14 are distributed within the area covered by the communication system 1. Each edge data center 14 is capable of communicating with communication equipment 18 equipped with an antenna 16. As shown in FIG. 1 , one edge data center 14 may be capable of communicating with several pieces of communication equipment 18. The communication equipment 18 may include a computer such as a server computer. The communication equipment 18 according to this embodiment performs wireless communication with a UE (User Equipment) 20 via the antenna 16. The communication equipment 18 equipped with the antenna 16 is provided with, for example, a radio unit (RU) (described later).
[0020] In this embodiment, the central data center 10, the regional data center 12, and the edge data center 14 each have a plurality of servers arranged therein.
[0021] In this embodiment, for example, the central data center 10, the regional data centers 12, and the edge data centers 14 are capable of communicating with each other. Furthermore, the central data centers 10, the regional data centers 12, and the edge data centers 14 are also capable of communicating with each other.
[0022] 2, the communication system 1 according to this embodiment includes a platform system 30, multiple radio access networks (RANs) 32, multiple core network systems 34, and multiple UEs 20. The core network systems 34, the RANs 32, and the UEs 20 cooperate with each other to realize a mobile communication network.
[0023] The RAN 32 is a computer system equipped with an antenna 16, which corresponds to an eNodeB (eNB) in a fourth-generation mobile communication system (hereinafter referred to as 4G) or a gNB (NR base station) in a fifth-generation mobile communication system (hereinafter referred to as 5G). The RAN 32 according to this embodiment is mainly implemented by a group of servers and communication equipment 18 arranged in an edge data center 14. Note that part of the RAN 32 (for example, a distributed unit (DU), a central unit (CU), a virtual distributed unit (vDU), and a virtual central unit (vCU)) may be implemented in the central data center 10 or the regional data center 12, rather than in the edge data center 14.
[0024] The core network system 34 is a system equivalent to an EPC (Evolved Packet Core) in 4G or a 5G Core (5GC) in 5G. The core network system 34 according to this embodiment is implemented mainly by a group of servers arranged in the central data center 10 and the regional data centers 12.
[0025] The platform system 30 according to this embodiment is configured, for example, on a cloud platform and includes a processor 30a, a storage unit 30b, and a communication unit 30c, as shown in FIG. 2 . The processor 30a is a program-controlled device such as a microprocessor that operates according to a program installed in the platform system 30. The storage unit 30b is, for example, a storage element such as a ROM or RAM, a solid-state drive (SSD), or a hard disk drive (HDD). The storage unit 30b stores programs executed by the processor 30a. The communication unit 30c is, for example, a communication interface such as a network interface controller (NIC) or a wireless local area network (LAN) module. Note that software-defined networking (SDN) may be implemented in the communication unit 30c. The communication unit 30c exchanges data with the RAN 32 and the core network system 34.
[0026] In this embodiment, the platform system 30 is implemented by a group of servers located in the central data center 10. Note that the platform system 30 may also be implemented by a group of servers located in the regional data centers 12.
[0027] In this embodiment, for example, in response to a purchase request for a network service (NS) by a purchaser, the requested network service is constructed in the RAN 32 or the core network system 34. Then, the constructed network service is provided to the purchaser.
[0028] For example, a purchaser such as an MVNO (Mobile Virtual Network Operator) is provided with network services such as voice communication services and data communication services. The voice communication services and data communication services provided by this embodiment are ultimately provided to customers (end users) of the purchaser (MVNO in the above example) who use the UE 20 shown in Figures 1 and 2. The end users can perform voice communication and data communication with other users via the RAN 32 and the core network system 34. The UE 20 of the end user can also access a data network such as the Internet via the RAN 32 and the core network system 34.
[0029] In addition, in this embodiment, an IoT (Internet of Things) service may be provided to an end user who uses a robot arm, a connected car, etc. In this case, for example, the end user who uses the robot arm, the connected car, etc. may become a purchaser of the network service according to this embodiment.
[0030] In this embodiment, a container-type virtualized application execution environment such as Docker (registered trademark) is installed on servers located in the central data center 10, the regional data centers 12, and the edge data center 14, allowing containers to be deployed and run on these servers. A cluster consisting of one or more containers generated by such virtualization technology may be constructed on these servers. For example, a Kubernetes cluster managed by a container management tool such as Kubernetes (registered trademark) may be constructed. Then, a processor on the constructed cluster may execute a container-type application.
[0031] In this embodiment, the network service provided to the purchaser is composed of one or more functional units (for example, network functions (NFs)). In this embodiment, the functional units are implemented as NFs realized by virtualization technology. NFs realized by virtualization technology are called VNFs (Virtualized Network Functions). It does not matter what virtualization technology is used for virtualization. For example, in this description, a CNF (Containerized Network Function) realized by container-type virtualization technology is also included in the VNF. In this embodiment, the network service is described as being implemented by one or more CNFs. Furthermore, the functional units according to this embodiment may correspond to network nodes.
[0032] Fig. 3 is a diagram illustrating an example of a network service in operation. The network service illustrated in Fig. 3 includes, as software elements, NFs such as a plurality of RUs 40, a plurality of DUs 42, a plurality of CUs 44 (CU-CPs (Central Unit - Control Plane) 44a and CU-UPs (Central Unit - User Plane) 44b), a plurality of AMFs (Access and Mobility Management Functions) 46, a plurality of SMFs (Session Management Functions) 48, and a plurality of UPFs (User Plane Functions) 50.
[0033] In the example of Figure 3, RU 40, DU 42, CU-CP 44a, AMF 46, and SMF 48 correspond to elements of the control plane (C-Plane), and RU 40, DU 42, CU-UP 44b, and UPF 50 correspond to elements of the user plane (U-Plane).
[0034] The network service may include other types of NF as software elements. The network service is implemented on computer resources (hardware elements) such as multiple servers.
[0035] In this embodiment, for example, a communication service in a certain area is provided by the network service shown in FIG.
[0036] In this embodiment, it is assumed that the multiple RUs 40, multiple DUs 42, multiple CU-UPs 44b, and multiple UPFs 50 shown in Figure 3 belong to one end-to-end network slice.
[0037] 4 is a diagram schematically illustrating an example of associations between elements established in the communication system 1 in this embodiment. The symbols M and N shown in FIG. 4 represent any integers equal to or greater than 1, and indicate the relationship between the numbers of elements connected by a link. When both ends of a link are a combination of M and N, the elements connected by the link have a many-to-many relationship, and when both ends of a link are a combination of 1 and N or a combination of 1 and M, the elements connected by the link have a one-to-many relationship.
[0038] As shown in FIG. 4, the network service (NS), network function (NF), CNFC (Containerized Network Function Component), pod, and container have a hierarchical structure.
[0039] The NS corresponds to, for example, a network service configured from a plurality of NFs. Here, the NS may correspond to, for example, a granular element such as 5GC, EPC, 5G RAN (gNB), or 4G RAN (eNB).
[0040] In 5G, NFs correspond to elements with granularity such as RU, DU, CU-CP, CU-UP, AMF, SMF, and UPF. In 4G, NFs correspond to elements with granularity such as MME (Mobility Management Entity), HSS (Home Subscriber Server), S-GW (Serving Gateway), vDU, and vCU. In this embodiment, for example, one NS includes one or more NFs. In other words, one or more NFs are under the control of one NS.
[0041] CNFC corresponds to a granularity element such as DU mgmt or DU Processing, for example. CNFC may be a microservice deployed on a server as one or more containers. For example, a certain CNFC may be a microservice that provides some of the functions of DU, CU-CP, CU-UP, etc. Also, a certain CNFC may be a microservice that provides some of the functions of UPF, AMF, SMF, etc. In this embodiment, for example, one NF includes one or more CNFCs. In other words, one or more CNFCs are under the control of one NF.
[0042] A pod is the smallest unit for managing a Docker container in Kubernetes. In this embodiment, for example, one CNFC includes one or more pods. In other words, one CNFC has one or more pods under its control.
[0043] In this embodiment, for example, one pod includes one or more containers, that is, one or more containers are subordinate to one pod.
[0044] Also, as shown in Figure 4, network slices (NSIs) and network slice subnet instances (NSSIs) have a hierarchical structure.
[0045] The NSI can also be considered an end-to-end virtual circuit spanning multiple domains (e.g., from the RAN 32 to the core network system 34). The NSI may be a slice for high-speed, large-capacity communication (e.g., for enhanced Mobile Broadband (eMBB)), a slice for high-reliability and low-latency communication (e.g., for Ultra-Reliable and Low Latency Communications (URLLC)), or a slice for connecting a large number of terminals (e.g., for massive Machine Type Communication (mMTC)). The NSSI can also be considered a virtual circuit of a single domain obtained by dividing the NSI. The NSSI may be a slice of the RAN domain, a slice of a transport domain such as a Mobile Back Haul (MBH) domain, or a slice of the core network domain.
[0046] In this embodiment, for example, one NSI includes one or more NSSIs. That is, one or more NSSIs are subordinate to one NSI. Note that in this embodiment, multiple NSIs may share the same NSSI.
[0047] Furthermore, as shown in FIG. 4, NSSIs and NSs generally have a many-to-many relationship.
[0048] Furthermore, in this embodiment, for example, one NF can belong to one or more network slices. Specifically, for example, one NF can be configured with NSSAI (Network Slice Selection Assistance Information) including one or more S-NSSAI (Sub Network Slice Selection Assist Information). Here, S-NSSAI is information associated with a network slice. Note that an NF does not necessarily have to belong to a network slice.
[0049] Fig. 5 is a functional block diagram showing an example of functions implemented in the platform system 30 according to this embodiment. Note that the platform system 30 according to this embodiment does not need to implement all of the functions shown in Fig. 5, and functions other than the functions shown in Fig. 5 may also be implemented.
[0050] As shown in FIG. 5 , the platform system 30 according to this embodiment functionally includes, for example, an operations support system (OSS) unit 60, an orchestration (E2EO: End-to-End-Orchestration) unit 62, a service catalog storage unit 64, a big data platform unit 66, a data bus unit 68, an AI (Artificial Intelligence) unit 70, a monitoring function unit 72, an SDN controller 74, a configuration management unit 76, a container management unit 78, and a repository unit 80. The OSS unit 60 includes an inventory database 82, a ticket management unit 84, a fault management unit 86, and a performance management unit 88. The E2EO unit 62 includes a policy manager unit 90, a slice manager unit 92, and a lifecycle management unit 94. These elements are implemented primarily using a processor 30 a, a storage unit 30 b, and a communication unit 30 c.
[0051] The functions shown in Figure 5 may be implemented by installing a program including instructions corresponding to the functions in a platform system 30, which is one or more computers, on the platform system 30 and having the processor 30a execute the program. This program may be supplied to the platform system 30 via a computer-readable information storage medium, such as an optical disk, a magnetic disk, a magnetic tape, a magneto-optical disk, or a flash memory, or via the Internet. The functions shown in Figure 5 may also be implemented using circuit blocks, memory, or other LSIs. Those skilled in the art will understand that the functions shown in Figure 5 can be realized in various forms, such as hardware alone, software alone, or a combination thereof.
[0052] The container management unit 78 manages the life cycle of a container, including processes related to the construction of the container, such as the deployment and configuration of the container.
[0053] Here, the platform system 30 according to the present embodiment may include a plurality of container management units 78. A container management tool such as Kubernetes and a package manager such as Helm may be installed in each of the plurality of container management units 78. Each of the plurality of container management units 78 may execute container construction, such as container deployment, for a server group (e.g., a Kubernetes cluster) associated with the corresponding container management unit 78.
[0054] The container management unit 78 does not need to be included in the platform system 30. The container management unit 78 may be provided, for example, in a server managed by the container management unit 78 (i.e., the RAN 32 or the core network system 34), or may be provided in another server that is annexed to the server managed by the container management unit 78.
[0055] In this embodiment, the repository unit 80 stores, for example, container images of containers included in a group of functional units (for example, a group of NFs) that realize a network service.
[0056] The inventory database 82 is a database that stores inventory information, which includes, for example, information about servers that are placed in the RAN 32 and the core network system 34 and that are managed by the platform system 30.
[0057] In this embodiment, inventory data is stored in the inventory database 82. The inventory data indicates the configuration of the elements included in the communication system 1 and the current status of the associations between the elements. The inventory data also indicates the status of resources managed by the platform system 30 (e.g., resource usage status). The inventory data may be physical inventory data or logical inventory data. Physical inventory data and logical inventory data will be described later.
[0058] Fig. 6 is a diagram showing an example of the data structure of physical inventory data. The physical inventory data shown in Fig. 6 is associated with one server. The physical inventory data shown in Fig. 6 includes, for example, a server ID, location data, building data, floor data, rack data, specification data, network data, an operating container ID list, a cluster ID, and the like.
[0059] The server ID included in the physical inventory data is, for example, an identifier of the server associated with the physical inventory data.
[0060] The location data included in the physical inventory data is, for example, data indicating the location (for example, the address of the location) of the server associated with the physical inventory data.
[0061] The building data included in the physical inventory data is, for example, data indicating the building (for example, the building name) in which the server associated with the physical inventory data is located.
[0062] The floor number data included in the physical inventory data is, for example, data indicating the floor number on which the server associated with the physical inventory data is located.
[0063] The rack data included in the physical inventory data is, for example, an identifier of the rack in which the server associated with the physical inventory data is located.
[0064] The specification data included in the physical inventory data is, for example, data indicating the specifications of the server associated with the physical inventory data, and the specification data indicates, for example, the number of cores, memory capacity, hard disk capacity, etc.
[0065] The network data included in the physical inventory data is, for example, data indicating information about the network of the server associated with the physical inventory data, and the network data indicates, for example, the NIC equipped in the server, the number of ports equipped in the NIC, the port ID of the port, etc.
[0066] The operating container ID list included in the physical inventory data is, for example, data that indicates information about one or more containers operating on a server associated with the physical inventory data, and the operating container ID list indicates, for example, a list of identifiers (container IDs) of instances of the containers.
[0067] The cluster ID included in the physical inventory data is, for example, an identifier of the cluster (for example, a Kubernetes cluster) to which the server associated with the physical inventory data belongs.
[0068] The logical inventory data includes topology data indicating the current state of associations between multiple elements included in the communication system 1, such as those shown in Figure 4. For example, the logical inventory data includes topology data including an identifier of a certain NS and identifiers of one or more NFs under the NS. Also, for example, the logical inventory data includes topology data including an identifier of a certain network slice and identifiers of one or more NFs belonging to the network slice.
[0069] The inventory data may also include data indicating the current status of geographical relationships and topological relationships between elements included in the communication system 1. As described above, the inventory data includes location data indicating the locations where the elements included in the communication system 1 are operating, i.e., the current locations of the elements included in the communication system 1. From this, it can be said that the inventory data indicates the current status of the geographical relationships between elements (e.g., the geographical proximity between elements).
[0070] The logical inventory data may also include NSI data indicating information about the network slice. The NSI data indicates attributes such as an identifier of an instance of the network slice and a type of the network slice. The logical inventory data may also include NSSI data indicating information about the network slice subnet. The NSSI data indicates attributes such as an identifier of an instance of the network slice subnet and a type of the network slice subnet.
[0071] The logical inventory data may also include NS data indicating information about NS. The NS data indicates attributes such as an NS instance identifier and an NS type, for example. The logical inventory data may also include NF data indicating information about NF. The NF data indicates attributes such as an NF instance identifier and an NF type, for example. The logical inventory data may also include CNFC data indicating information about CNFC. The CNFC data indicates attributes such as an instance identifier and a CNFC type, for example. The logical inventory data may also include pod data indicating information about pods included in the CNFC. The pod data indicates attributes such as a pod instance identifier and a pod type, for example. The logical inventory data may also include container data indicating information about containers included in the pod. The container data indicates attributes such as a container ID of a container instance and a container type, for example.
[0072] The container ID of the container data included in the logical inventory data and the container ID included in the operating container ID list included in the physical inventory data associate a container instance with the server on which the container instance is running.
[0073] Furthermore, the logical inventory data may include data indicating various attributes such as a host name and an IP address. For example, the container data may include data indicating an IP address of a container corresponding to the container data. For example, the NF data may include data indicating an IP address and a host name of the NF indicated by the NF data.
[0074] The logical inventory data may also include data indicating an NSSAI, including one or more S-NSSAIs, that is set in each NF.
[0075] The inventory database 82 is also able to grasp the resource status as needed in cooperation with the container management unit 78. The inventory database 82 then updates the inventory data stored therein as needed based on the latest resource status.
[0076] In addition, in response to actions being performed, such as constructing a new element included in the communication system 1, changing the configuration of an element included in the communication system 1, scaling an element included in the communication system 1, or replacing an element included in the communication system 1, the inventory database 82 updates the inventory data stored in the inventory database 82.
[0077] The service catalog storage unit 64 stores service catalog data. The service catalog data may include, for example, service template data indicating logic used by the life cycle management unit 94. This service template data includes information necessary for building a network service. For example, the service template data includes information defining NS, NF, and CNFC, and information indicating the correspondence between NS, NF, and CNFC. Furthermore, for example, the service template data includes a workflow script for building a network service.
[0078] An example of service template data is an NSD (NS Descriptor). The NSD is associated with a network service and indicates the types of multiple functional units (e.g., multiple CNFs) included in the network service. The NSD may also indicate the number of each type of functional unit, such as a CNF, included in the network service. The NSD may also indicate the file name of a CNFD (described later) related to the CNF included in the network service.
[0079] An example of service template data is a CNF Descriptor (CNFD). The CNFD may indicate computer resources (e.g., a CPU, memory, hard disk, etc.) required by the CNF. For example, the CNFD may indicate, for each of multiple containers included in the CNF, the computer resources (e.g., a CPU, memory, hard disk, etc.) required by the container.
[0080] The service catalog data may also include information about thresholds (for example, anomaly detection thresholds) that are used by the policy manager 90 to compare with the calculated performance index values. The performance index values will be described later.
[0081] The service catalog data may also include, for example, slice template data, which includes information necessary to perform instantiation of a network slice, including, for example, logic utilized by the slice manager unit 92.
[0082] The slice template data includes information on the "Generic Network Slice Template" defined by the GSM Association (GSMA) ("GSM" is a registered trademark). Specifically, the slice template data includes network slice template data (NST), network slice subnet template data (NSST), and network service template data. The slice template data also includes information indicating the hierarchical structure of these elements, as shown in FIG. 4.
[0083] In this embodiment, for example, the life cycle management unit 94 constructs a new network service in response to a purchase request for an NS from a purchaser.
[0084] For example, in response to a purchase request, the lifecycle management unit 94 may execute a workflow script associated with the network service to be purchased. By executing this workflow script, the lifecycle management unit 94 may instruct the container management unit 78 to deploy a container included in the new network service to be purchased. The container management unit 78 may then obtain a container image of the container from the repository unit 80 and deploy the container corresponding to the container image to a server.
[0085] In addition, in this embodiment, the life cycle management unit 94 executes, for example, scaling and replacement of elements included in the communication system 1. Here, the life cycle management unit 94 may output a container deployment instruction or deletion instruction to the container management unit 78. Then, the container management unit 78 may execute processing such as container deployment or container deletion in accordance with the instruction. In this embodiment, the life cycle management unit 94 is capable of executing scaling and replacement that cannot be handled by a tool such as Kubernetes in the container management unit 78.
[0086] Furthermore, the life cycle management unit 94 may output an instruction to create a communication path to the SDN controller 74. For example, the life cycle management unit 94 presents two IP addresses at both ends of the communication path to be created to the SDN controller 74, and the SDN controller 74 creates a communication path connecting these two IP addresses. The created communication path may be managed in association with these two IP addresses.
[0087] Furthermore, the life cycle management unit 94 may output to the SDN controller 74 an instruction to create a communication path between the two IP addresses that is associated with the two IP addresses.
[0088] In this embodiment, for example, the slice manager unit 92 performs instantiation of a network slice. In this embodiment, for example, the slice manager unit 92 performs instantiation of a network slice by executing logic indicated by a slice template stored in the service catalog storage unit 64.
[0089] The slice manager unit 92 is configured to include the functions of the NSMF (Network Slice Management Function) and the NSSMF (Network Slice Sub-network Management Function), for example, as described in the specification "TS28 533" of the 3GPP (registered trademark) (Third Generation Partnership Project). The NSMF is a function that generates and manages network slices and provides NSI management services. The NSSMF is a function that generates and manages network slice subnets that constitute part of the network slice and provides NSSI management services.
[0090] Here, the slice manager unit 92 may output a configuration management instruction related to the instantiation of the network slice to the configuration management unit 76. Then, the configuration management unit 76 may perform configuration management such as setting in accordance with the configuration management instruction.
[0091] The slice manager unit 92 may also present two IP addresses to the SDN controller 74 and output an instruction to create a communication path between these two IP addresses.
[0092] In this embodiment, the configuration management unit 76 performs configuration management such as setting of element groups such as NFs in accordance with configuration management instructions received from the life cycle management unit 94 and the slice manager unit 92, for example.
[0093] In this embodiment, the SDN controller 74 creates a communication path between two IP addresses associated with a communication path creation instruction received from, for example, the life cycle management unit 94 or the slice manager unit 92. The SDN controller 74 may create a communication path between two IP addresses using a known path calculation method such as Flex Algo.
[0094] Here, for example, the SDN controller 74 may use a segment routing technology (for example, SRv6 (Segment Routing IPv6)) to construct NSIs and NSSIs for aggregation routers, servers, and the like present along the communication paths. Furthermore, the SDN controller 74 may generate NSIs and NSSIs across multiple NFs to be configured by issuing commands to configure a common VLAN (Virtual Local Area Network) for multiple NFs to be configured, and commands to assign the bandwidth and priority indicated in the configuration information to the VLAN.
[0095] In addition, the SDN controller 74 may perform operations such as changing the maximum bandwidth available for communication between two IP addresses without constructing a network slice.
[0096] The platform system 30 according to this embodiment may include multiple SDN controllers 74. Each of the multiple SDN controllers 74 may execute processing such as creating a communication path for a group of network devices such as an AG associated with the SDN controller 74.
[0097] In this embodiment, for example, the monitoring function unit 72 monitors the group of elements included in the communication system 1 in accordance with a given management policy. Here, the monitoring function unit 72 may monitor the group of elements in accordance with a monitoring policy specified by a purchaser when purchasing a network service, for example.
[0098] In this embodiment, the monitoring function unit 72 performs monitoring at various levels, such as the slice level, the NS level, the NF level, the CNFC level, and the hardware level of a server or the like.
[0099] For example, in order to perform monitoring at the various levels described above, the monitoring function unit 72 may set a module that outputs metric data in hardware such as a server or in a software element included in the communication system 1. Here, for example, an NF may output metric data indicating metrics that are measurable (identifiable) in the NF to the monitoring function unit 72. Also, a server may output metric data indicating metrics related to hardware that is measurable (identifiable) in the server to the monitoring function unit 72.
[0100] Furthermore, for example, the monitoring function unit 72 may deploy a sidecar container on the server that aggregates metric data indicating metrics output from multiple containers on a CNFC (microservice) basis. This sidecar container may include an agent called an exporter. The monitoring function unit 72 may repeatedly execute, at a given monitoring interval, a process of acquiring metric data aggregated on a microservice basis from the sidecar container using a mechanism of a monitoring tool such as Prometheus, which can monitor container management tools such as Kubernetes.
[0101] The monitoring function unit 72 may monitor performance indicator values for performance indicators described in, for example, “TS 28.552, Management and orchestration; 5G performance measurements” or “TS 28.554, Management and orchestration; 5G end to end Key Performance Indicators (KPI).” Then, the monitoring function unit 72 may acquire metric data indicating the monitored performance indicator values.
[0102] In this embodiment, the monitoring function unit 72 performs a process (enrichment) of aggregating metric data, for example, in a predetermined aggregation unit, thereby generating performance index value data indicating the performance index values of the elements included in the communication system 1 in that aggregation unit.
[0103] For example, for one gNB, performance index value data for the gNB is generated by aggregating metric data indicating the metrics of elements (e.g., network nodes such as DU42 and CU44) under the control of the gNB. In this way, performance index value data indicating communication performance in the area covered by the gNB is generated. Here, for example, performance index value data indicating multiple types of communication performance such as traffic volume (throughput) and latency may be generated for each gNB. Note that the communication performance indicated by the performance index value data is not limited to traffic volume and latency.
[0104] Then, the monitoring function unit 72 outputs the performance index value data generated by the above-mentioned enrichment to the data bus unit 68.
[0105] In this embodiment, for example, the data bus unit 68 receives performance index value data output from the monitoring function unit 72. Then, based on the received one or more pieces of performance index value data, the data bus unit 68 generates a performance index value file including the one or more pieces of performance index value data. Then, the data bus unit 68 outputs the generated performance index value file to the big data platform unit 66.
[0106] In this embodiment, the monitoring function unit 72 identifies a stability evaluation value indicating the stability of each application by, for example, executing a process (enrichment) of aggregating metric data related to the application. Then, the monitoring function unit 72 generates stability evaluation value data indicating the identified stability evaluation value.
[0107] Then, the monitoring function unit 72 outputs the generated stability evaluation value data to the data bus unit 68 .
[0108] In this embodiment, the data bus unit 68 receives, for example, stability evaluation value data output from the monitoring function unit 72 .
[0109] In addition, elements such as network slices, NS, NF, CNFC, etc. included in the communication system 1, and hardware such as servers, notify the monitoring function unit 72 of various alerts (for example, notification of an alert triggered by the occurrence of a failure).
[0110] Then, for example, when the monitoring function unit 72 receives the above-mentioned alert notification, it outputs alert message data indicating the notification to the data bus unit 68. Then, the data bus unit 68 generates an alert file in which alert message data indicating one or more notifications are compiled into a single file, and outputs the alert file to the big data platform unit 66.
[0111] In this embodiment, the big data platform unit 66 accumulates, for example, performance index value files and alert files output from the data bus unit 68 .
[0112] In this embodiment, for example, a plurality of trained machine learning models are stored in advance in the AI unit 70. The AI unit 70 uses the various machine learning models stored in the AI unit 70 to perform estimation processing such as future prediction processing of the usage status and service quality of the communication system 1. The AI unit 70 may generate estimation result data indicating the results of the estimation processing.
[0113] The AI unit 70 may perform estimation processing based on the files stored in the big data platform unit 66 and the above-mentioned machine learning model. This estimation processing is suitable for low-frequency prediction of long-term trends.
[0114] The AI unit 70 is also capable of acquiring performance index value data stored in the data bus unit 68. The AI unit 70 may perform estimation processing based on the performance index value data stored in the data bus unit 68 and the above-described machine learning model. This estimation processing is suitable for performing short-term predictions frequently.
[0115] In this embodiment, for example, the performance management unit 88 calculates a performance index value (e.g., KPI) based on metrics indicated by multiple metric data. The performance management unit 88 may calculate a performance index value that is an overall evaluation of multiple types of metrics (e.g., a performance index value related to an end-to-end network slice) that cannot be calculated from a single metric data. The performance management unit 88 may generate overall performance index value data that indicates the performance index value that is the overall evaluation.
[0116] The performance management unit 88 may acquire the above-mentioned performance index value file from the big data platform unit 66. The performance management unit 88 may also acquire estimation result data from the AI unit 70. Then, performance index values such as KPIs may be calculated based on at least one of the performance index value file and the estimation result data. The performance management unit 88 may also directly acquire metric data from the monitoring function unit 72. Then, performance index values such as KPIs may be calculated based on the metric data.
[0117] In this embodiment, the fault management unit 86 detects the occurrence of a fault in the communication system 1 based on, for example, at least one of the above-mentioned metric data, the above-mentioned alert notification, the above-mentioned estimation result data, and the above-mentioned overall performance index value data. The fault management unit 86 may detect the occurrence of a fault that cannot be detected from a single piece of metric data or a single alert notification, for example, based on a predetermined logic. The fault management unit 86 may generate detected fault data that indicates the detected fault.
[0118] The fault management unit 86 may obtain metric data and alert notifications directly from the monitoring function unit 72. The fault management unit 86 may also obtain performance index value files and alert files from the big data platform unit 66. The fault management unit 86 may also obtain alert message data from the data bus unit 68.
[0119] In this embodiment, the policy manager unit 90 executes a predetermined judgment process based on, for example, at least one of the above-mentioned metric data, the above-mentioned performance index value data, the above-mentioned stability evaluation value data, the above-mentioned alert message data, the above-mentioned performance index value file, the above-mentioned alert file, the above-mentioned estimation result data, the above-mentioned overall performance index value data, and the above-mentioned detected fault data.
[0120] The policy manager unit 90 may then execute an action according to the result of the determination process. For example, the policy manager unit 90 may output an instruction to construct a network slice to the slice manager unit 92. The policy manager unit 90 may also output an instruction to scale or replace an element to the life cycle management unit 94 according to the result of the determination process.
[0121] The policy manager unit 90 according to this embodiment is capable of acquiring performance index value data stored in the data bus unit 68. The policy manager unit 90 may then execute a predetermined determination process based on the performance index value data acquired from the data bus unit 68. The policy manager unit 90 may also execute a predetermined determination process based on alert message data stored in the data bus unit 68.
[0122] Furthermore, the policy manager unit 90 according to this embodiment is capable of acquiring stability evaluation value data stored in the data bus unit 68. The policy manager unit 90 may then execute a predetermined determination process based on the stability evaluation value data acquired from the data bus unit 68. For example, the policy manager unit 90 may determine whether an application is unstable based on the stability evaluation value data indicating the stability of the application.
[0123] In this embodiment, for example, the ticket management unit 84 generates a ticket indicating the content to be notified to the administrator of the communication system 1. The ticket management unit 84 may generate a ticket indicating the content of the occurred fault data. The ticket management unit 84 may also generate a ticket indicating the values of performance index value data, stability evaluation value data, or metric data. The ticket management unit 84 may also generate a ticket indicating the determination result by the policy manager unit 90.
[0124] Then, the ticket management unit 84 notifies the administrator of the communication system 1 of the generated ticket. For example, the ticket management unit 84 may send an email with the generated ticket attached to the email address of the administrator of the communication system 1.
[0125] As described above, in this embodiment, the policy manager unit 90 determines whether an application is unstable based on a stability evaluation value indicating the stability of the application. If the application is determined to be unstable, the policy manager unit 90 estimates the cause of the application's instability. For example, the policy manager unit 90 estimates whether the cause of the application's instability lies in the application itself or in the hardware resources on which a process included in the application is running.
[0126] The process may be, for example, an execution unit (e.g., a pod) of the application in a container-type virtualized application execution environment.
[0127] The application may also be a network function (e.g., DU 42, CU-CP 44a, CU-UP 44b, AMF 46, SMF 48, UPF 50, etc.).
[0128] The process of estimating the cause of application instability will be further described below.
[0129] In this embodiment, for example, as described above, the monitoring function unit 72 calculates a stability index value indicating the stability of each of the multiple applications included in the communication system 1. These applications include multiple types of processes. For each type, multiple processes of that type run, resulting in the entire application running. Furthermore, for each type, the processes of that type run in a distributed manner across multiple hardware resources.
[0130] FIG. 7 is a diagram illustrating an example of a situation in which processes included in a plurality of applications are distributed and run on a plurality of hardware resources.
[0131] The example of FIG. 7 shows a situation in which four applications with identifiers AP1, AP2, AP3, and AP4 are running.
[0132] In this embodiment, for each type of application, a hardware resource on which the application of that type can run is predetermined. In the following description, the hardware resource is assumed to be a server, but the hardware resource does not have to be a server and may be, for example, a node.
[0133] Hereinafter, a hardware resource on which a certain type of application can run will be referred to as a tenant corresponding to that application.
[0134] 7 shows four servers with identifiers S1, S2, S3, and S4, respectively. These four servers belong to one cluster (for example, a Kubernetes cluster).
[0135] 7 are assumed to be of different types. The tenant corresponding to the application with the identifier AP1 includes servers with identifiers S1, S2, and S3. The tenant corresponding to the application with the identifier AP2 includes servers with identifiers S3 and S4. The tenant corresponding to the application with the identifier AP3 includes servers with identifiers S1, S2, S3, and S4. The tenant corresponding to the application with the identifier AP4 includes servers with identifiers S1 and S4.
[0136] 7 corresponds to one process (e.g., a pod). The numbers shown in the rounded rectangles are identifiers associated with the process types. That is, rounded rectangles with the same numbers correspond to processes of the same type.
[0137] 7, an application with an identifier AP1 includes three types of processes with identifiers 1, 2, and 3. The three processes of the type with identifier 1 are running on servers with identifiers S1, S2, and S3, respectively. Three processes of the type with identifier 2 are running on servers with identifiers S1, S2, and S3, respectively. Two processes of the type with identifier 3 are running on servers with identifiers S1 and S2, respectively.
[0138] Furthermore, an application with an identifier of AP2 includes four types of processes with identifiers of 4, 5, 6, and 7. Two processes of the type with identifier 4 are running on servers with identifiers S3 and S4, respectively. One process of the type with identifier 5 is running on a server with identifier S3. Two processes of the type with identifier 6 are running on servers with identifiers S3 and S4, respectively. Two processes of the type with identifier 7 are running on servers with identifiers S3 and S4, respectively.
[0139] Furthermore, an application with an identifier AP3 includes three types of processes with identifiers 8, 9, and 10. Four processes of the type with identifier 8 are running on servers with identifiers S1, S2, S3, and S4, respectively. Three processes of the type with identifier 9 are running on servers with identifiers S1, S3, and S4, respectively. Three processes of the type with identifier 10 are running on servers with identifiers S2, S3, and S4, respectively.
[0140] Furthermore, an application with an identifier AP4 includes three types of processes with identifiers 11, 12, and 13. Two processes of the type with identifier 11 are running on servers with identifiers S1 and S4, respectively. Two processes of the type with identifier 12 are running on servers with identifiers S1 and S4, respectively. One process of the type with identifier 13 is running on the server with identifier S4.
[0141] In this embodiment, for example, the container management unit 78 controls each type of process so that it runs in a distributed manner across as many hardware resources as possible.
[0142] In this embodiment, for example, the monitoring function unit 72 acquires, for each type of process, a value (metric) indicating the stability of the process of that type. Here, for example, metrics such as a value indicating the state of the process (e.g., kube_pod_status_ready), a start time of the process (e.g., kube_pod_start_time), the length of time the process performed input / output (e.g., container_fs_io_time_seconds_total), the number of dropped transmit packets of the process (e.g., container_network_transmit_packets_dropped_total), and the number of dropped receive packets of the process (e.g., container_network_receive_packets_dropped_total) may be acquired.
[0143] The monitoring function unit 72 then calculates a weighted sum of the acquired metrics using weights associated with the respective types as a stability evaluation value indicating the stability of the process of that type. For example, a weight for each type of metric may be predetermined for each process type. The weighted sum of the acquired metrics using the predetermined weights may then be calculated as a stability evaluation value indicating the stability of the process of that type. Hereinafter, a stability evaluation value indicating the stability of a process will be referred to as a process stability evaluation value. For example, in the example of FIG. 7 , a process stability evaluation value is calculated for each of the process types with identifiers 1 to 13.
[0144] The monitoring function unit 72 then identifies a stability evaluation value indicating the stability of the application based on the process stability evaluation value of the process related to each type of process included in the application, which is acquired for each type of process. Hereinafter, the stability evaluation value indicating the stability of an application will be referred to as the application stability evaluation value. For example, the monitoring function unit 72 calculates the application stability evaluation value of each application based on the process stability evaluation values calculated for the processes included in the application.
[0145] Here, the application stability evaluation value may be determined based on at least one of the following: the state of a process included in the application; the lifetime of a process included in the application; the length of time that a process included in the application has performed input / output; or the number of packet drops of a process included in the application. Here, for example, the lifetime of a process can be determined based on a value indicating the start time of the process.
[0146] Furthermore, the monitoring function unit 72 may calculate an application stability evaluation value indicating the stability of an application in accordance with a rule associated with the type of application. For example, a formula may be defined in advance for each type of application. Then, the application stability evaluation value of the application may be calculated by applying the process stability evaluation value of the process related to each type, which is obtained for each type of process included in the application, to the formula.
[0147] For example, an application stability evaluation value for an application with an identifier AP1 is calculated based on the process stability evaluation values of processes with identifiers 1 to 3. Furthermore, an application stability evaluation value for an application with an identifier AP2 is calculated based on the process stability evaluation values of processes with identifiers 4 to 7. Furthermore, an application stability evaluation value for an application with an identifier AP3 is calculated based on the process stability evaluation values of processes with identifiers 8 to 10. Furthermore, an application stability evaluation value for an application with an identifier AP4 is calculated based on the process stability evaluation values of processes with identifiers 11 to 13.
[0148] The monitoring function unit 72 then generates stability evaluation value data for each of the multiple applications, indicating the application stability evaluation value calculated for that application, and outputs the generated stability evaluation value data to the data bus unit 68. In this embodiment, for example, the monitoring function unit 72 generates stability evaluation value data at predetermined time intervals based on the latest situation. The monitoring function unit 72 then outputs the stability evaluation value data to the data bus unit 68 every time the stability evaluation value data is generated.
[0149] Then, in response to the stability evaluation value data being output to the data bus unit 68, the policy manager unit 90 acquires the stability evaluation value data. The policy manager unit 90 then identifies the application stability evaluation value indicated by the acquired stability evaluation value data. In this way, the policy manager unit 90 identifies, for each of a plurality of applications, a stability evaluation value indicating the stability of the application. Furthermore, as described above, the processes included in these applications are distributed and run on a plurality of hardware resources.
[0150] The policy manager unit 90 then determines whether each of the multiple applications is unstable based on a stability evaluation value that indicates the stability of the application. For example, the more unstable the application, the smaller the application stability evaluation value. In this case, the policy manager unit 90 determines that the application is unstable when, for example, the application stability evaluation value is smaller than a threshold value associated with the type of application.
[0151] For example, suppose that the policy manager unit 90 determines that a first application (e.g., an application with an identifier AP1) is unstable. In this case, the policy manager unit 90 identifies a stability evaluation value indicating the stability of a second application having at least one process running on a hardware resource on which at least one process included in the first application is running.
[0152] In the example of FIG. 7, the servers on which processes included in the first application are running are three servers with identifiers S1, S2, and S3.
[0153] In addition to the process included in the application with the identifier AP1, a process included in the application with the identifier AP3 and a process included in the application with the identifier AP4 are running on the server with the identifier S1.
[0154] Furthermore, in the server with the identifier S2, in addition to the process included in the application with the identifier AP1, a process included in the application with the identifier AP3 is running.
[0155] Furthermore, on the server with identifier S3, in addition to the process included in the application with identifier AP1, a process included in the application with identifier AP2 and a process included in the application with identifier AP3 are running.
[0156] Therefore, in this case, the three applications with identifiers AP2, AP3, and AP4 correspond to the second application described above. In this way, there may be a plurality of second applications.
[0157] Therefore, in this case, the policy manager unit 90 identifies the application stability evaluation value for each of the three applications with identifiers AP2, AP3, and AP4.
[0158] When the policy manager unit 90 determines that the first application is unstable based on a stability evaluation value indicating the stability of the first application, it estimates, based on a stability evaluation value indicating the stability of the second application, whether the instability of the first application is caused by the first application itself or by the hardware resources on which the processes included in the first application are running.
[0159] Here, if the policy manager unit 90 determines that the second application is not unstable based on a stability evaluation value indicating the stability of the second application, it may infer that the first application is the cause of the instability of the first application.
[0160] For example, if it is determined that an application with an identifier AP2 is not unstable, it may be assumed that the instability of the application with an identifier AP1 is caused by that application. Alternatively, if it is determined that an application with an identifier AP3 is not unstable, it may be assumed that the instability of the application with an identifier AP1 is caused by that application. Alternatively, if it is determined that an application with an identifier AP4 is not unstable, it may be assumed that the instability of the application with an identifier AP1 is caused by that application.
[0161] In addition, when the policy manager unit 90 determines that the second application is unstable based on a stability evaluation value indicating the stability of the second application, it may infer that the instability of the first application is caused by the hardware resources on which the processes included in the first application are running.
[0162] For example, if an application with an identifier AP2 is determined to be unstable, it may be inferred that the instability of the application with an identifier AP1 is caused by the server with an identifier S3. Alternatively, if an application with an identifier AP3 is determined to be unstable, it may be inferred that the instability of the application with an identifier AP1 is caused by the server with an identifier S1, S2, or S3. Alternatively, if an application with an identifier AP4 is determined to be not unstable, it may be inferred that the instability of the application with an identifier AP1 is caused by the server with an identifier S1.
[0163] In addition, the policy manager unit 90 may estimate whether the instability of a first application is caused by the first application or the hardware resource, based on at least one of the number of applications that are determined to be unstable or the number of applications that are determined to be not unstable, among multiple applications running on any of the hardware resources on which a process included in the first application is running.
[0164] In this case, the policy manager unit 90 may infer that the reason the first application is unstable is that the hardware resource has a predetermined number of applications determined to be unstable, which is two or more.
[0165] For example, assume that the predetermined number is 3. In this case, if it is determined that three applications with identifiers AP1, AP3, and AP4 are unstable, it may be inferred that the instability of the first application is caused by the server with identifier S1. Also, if it is determined that three applications with identifiers AP1, AP2, and AP3 are unstable, it may be inferred that the instability of the first application is caused by the server with identifier S3. And, if none of the above cases apply, it may be inferred that the instability of the first application is caused by that application.
[0166] Alternatively, the policy manager unit 90 may estimate that the instability of the first application is caused by a hardware resource where the ratio of the number of applications determined to be unstable to the number of applications running on the hardware resource is greater than or equal to a predetermined value.
[0167] For example, assume that the predetermined value is 60%. In this case, if 60% or more of the applications running on a server with an identifier S1 are determined to be unstable, it may be inferred that the instability of the first application is caused by the server with an identifier S1. Alternatively, if 60% or more of the applications running on a server with an identifier S2 are determined to be unstable, it may be inferred that the instability of the first application is caused by the server with an identifier S2. Alternatively, if 60% or more of the applications running on a server with an identifier S3 are determined to be unstable, it may be inferred that the instability of the first application is caused by the server with an identifier S3. Then, if neither of the above cases occurs, it may be inferred that the instability of the first application is caused by that application.
[0168] In addition, if the policy manager unit 90 determines that all applications running on any of the hardware resources on which the processes included in the first application are running are unstable, it may infer that the instability of the first application is caused by that hardware resource.
[0169] For example, if all applications (three applications with identifiers AP1, AP3, and AP4) whose processes are running on a server with an identifier S1 are determined to be unstable, it may be inferred that the instability of the first application is caused by the server with an identifier S1. Alternatively, if all applications (two applications with identifiers AP1 and AP3) whose processes are running on a server with an identifier S2 are determined to be unstable, it may be inferred that the instability of the first application is caused by the server with an identifier S2. Alternatively, if all applications (three applications with identifiers AP1, AP2, and AP3) whose processes are running on a server with an identifier S3 are determined to be unstable, it may be inferred that the instability of the first application is caused by the server with an identifier S3. If neither of the above cases is true, it may be inferred that the instability of the first application is caused by that application.
[0170] In this case, if it is determined that the three applications with identifiers AP1, AP3, and AP4 are unstable, it may be presumed that the instability of the first application is caused by the server with identifier S1 or S2. Also, if it is determined that the three applications with identifiers AP1, AP2, and AP3 are unstable, it may be presumed that the instability of the first application is caused by the server with identifier S2 or S3.
[0171] Furthermore, if it is determined that three applications with identifiers AP1, AP3, and AP4 are unstable, it may be presumed that the instability of the first application is caused by the server with identifier S1 and the server with identifier S2. Furthermore, if it is determined that three applications with identifiers AP1, AP2, and AP3 are unstable, it may be presumed that the instability of the first application is caused by the server with identifier S2 and the server with identifier S3.
[0172] In this way, it is estimated whether the cause of instability in the first application lies in the first application itself or in the hardware resources on which the processes included in the first application are running.
[0173] The policy manager unit 90 then executes an action according to the estimated cause.
[0174] Here, the policy manager unit 90 may execute replacement of the first application. For example, if it is estimated that the instability of an application with an identifier AP1 is caused by that application, the application may be replaced with a server in another cluster or with another server in the same cluster. Here, for example, the tenant settings of the application may be changed.
[0175] The policy manager unit 90 may also execute separation of hardware resources from a cluster generated by virtualization technology. For example, if it is estimated that the cause of instability in an application with an identifier AP1 is a server with an identifier S1, the server may be separated from the cluster to which it belongs.
[0176] For example, as shown in FIG. 8 , a new server with an identifier S5 may be added to a cluster to which servers with identifiers S1 to S4 belong. Then, the tenant settings for AP1, AP3, and AP4 may be changed. For example, a server with an identifier S1 may be excluded from the tenant, and a server with an identifier S5 may be added to the tenant. Then, the server with an identifier S1 may be separated from the cluster to which it belongs. In this way, as shown in FIG. 8 , the container management unit 78 runs a process on the server with an identifier S5 as needed.
[0177] In this embodiment, when an application is determined to be unstable based on the application stability evaluation value aggregated for multiple processes, the cause of the instability of the application is estimated based on the application stability evaluation values of other applications.
[0178] Therefore, even when the stability evaluation value of an application is identified based on the application stability evaluation values collected for a plurality of processes, it is possible to accurately estimate the cause of the instability of the application.
[0179] In addition, in this embodiment, when an application is determined to be unstable, the policy manager unit 90 does not need to estimate the cause of the instability of the application as described above. In addition, in this embodiment, when an application is determined to be unstable, the process replacement may be performed in stages.
[0180] The gradual replacement of the process is further explained below.
[0181] In the following description, a group of hardware resources on which processes included in an application are distributed and running will be referred to as a group of running resources.
[0182] Furthermore, the hardware resources included in the group of operating resources in the initial state are referred to as replacement candidate resources. Here, as shown in Figure 7, for an application with identifier AP1, the servers with identifiers S1, S2, and S3 are replacement candidate resources, respectively. In other words, in the initial state, the application with identifier AP1 runs its processes distributed across a group of operating resources that includes multiple replacement candidate resources.
[0183] As described above, the policy manager unit 90 determines whether or not the application having the identifier AP1 is unstable based on the stability evaluation value data indicating the stability of the application.
[0184] Then, it is assumed that the application with the identifier AP1, whose processes are distributed among the group of operating resources, is determined to be unstable based on a stability evaluation value indicating the stability of the application.
[0185] In this case, the policy manager 90 may execute a configuration change to remove at least one replacement candidate resource from the group of operating resources and add a new resource to the group of operating resources. Here, for example, the configuration change of the tenant described above may be executed.
[0186] 9, one replacement candidate resource (e.g., a server with an identifier S3) may be removed from the group of operating resources for an application with an identifier AP1, and a new hardware resource (e.g., a server with an identifier S5) may be added to the group of operating resources.
[0187] 9, the container management unit 78 controls the processes included in the application with the identifier AP1 to run in a distributed manner across as many hardware resources as possible. That is, in this case, the processes included in the application with the identifier AP1 are controlled to run in a distributed manner across the servers with the identifiers S1, S2, and S5.
[0188] In this embodiment, each time the above-mentioned setting change is executed, the policy manager unit 90 may determine whether the application is unstable based on a stability evaluation value that indicates the stability of the application whose processes are running distributed among the group of operating resources after the setting change is executed.
[0189] In addition, the policy manager unit 90 may determine whether the application is unstable based on a stability evaluation value that indicates the stability of the application whose processes are running distributed across the group of operating resources after the setting change is executed, depending on whether a predetermined time (e.g., 15 minutes) has elapsed since the setting change was executed.
[0190] For example, whether an application with identifier AP1 is unstable may be determined based on a stability evaluation value indicating the stability of the application, whose processes are running in a distributed manner on servers with identifiers S1, S2, and S5.
[0191] Then, the policy manager unit 90 may execute the setting change again if it determines that the application is unstable based on a stability evaluation value indicating the stability of the application whose processes are running distributed across the group of operating resources after the above-mentioned setting change is executed.
[0192] For example, suppose a process included in an application with an identifier AP1 is distributed and running on servers with identifiers S1, S2, and S5. In this situation, suppose the application is determined to be unstable. In this case, as shown in FIG. 10 , one replacement candidate resource (e.g., a server with an identifier S2) may be excluded from the group of running resources for the application with an identifier AP1. Then, a new hardware resource (e.g., a server with an identifier S6) may be added to the group of running resources.
[0193] As a result, as shown in FIG. 10, the container management unit 78 controls the processes included in the application with identifier AP1 to run in a distributed manner on servers with identifiers S1, S5, and S6.
[0194] Whether an application having an identifier AP1 is unstable may be determined based on a stability evaluation value indicating the stability of the application having an identifier AP1, whose processes are distributed among servers having identifiers S1, S5, and S6. Here, whether an application is unstable may be determined based on the elapse of a predetermined time (e.g., 15 minutes) after the server having an identifier S2 is removed from the group of operating resources for the application having identifier AP1 and the server having an identifier S6 is added to the group of operating resources.
[0195] Then, suppose that an application is determined to be unstable based on a stability evaluation value indicating the stability of the application's processes distributed across the group of operating resources after the configuration change is executed. In this case, as shown in Figure 11, one replacement candidate resource (e.g., a server with an identifier S1) may be removed from the group of operating resources for an application with an identifier AP1. Then, a new hardware resource (e.g., a server with an identifier S7) may be added to the group of operating resources.
[0196] As shown in FIG. 11, the container management unit 78 may control the processes included in the application with the identifier AP1 to run in a distributed manner on servers with the identifiers S5, S6, and S7.
[0197] As described above, in this embodiment, the policy manager unit 90 may repeat the above-mentioned setting changes until the application is determined to be stable based on a stability evaluation value indicating the stability of the application whose processes are running distributed across the group of operating resources, or until all replacement candidate resources are excluded from the group of operating resources.
[0198] In this embodiment, the policy manager unit 90 may determine a replacement candidate resource to be excluded from the group of operating resources based on a stability evaluation value indicating the stability of other applications running on each of the replacement candidate resources included in the group of operating resources.The policy manager unit 90 may then exclude the determined replacement candidate resource from the group of operating resources.
[0199] For example, a replacement candidate resource having the largest number of applications determined to be unstable may be determined as the replacement candidate resource to be excluded from the group of operating resources.
[0200] Alternatively, a replacement candidate resource having the highest ratio of the number of applications determined to be unstable to the number of running applications may be determined as the replacement candidate resource to be excluded from the group of running resources.
[0201] 7, suppose that an application with an identifier AP1 is determined to be unstable. Then, suppose that an application with an identifier AP4 is determined to be unstable, and applications with identifiers AP2 and AP3 are determined to be stable. In this case, the server with an identifier S1 may be excluded from the group of operating resources for the application with identifier AP1.
[0202] In this embodiment, as described above, if a configuration change is performed to remove at least one replacement candidate resource from the group of operating resources and add a new resource to the group of operating resources, and the application is determined to be unstable based on a stability evaluation value indicating the stability of the application whose processes are running in a distributed manner across the group of operating resources after the configuration change is performed, the configuration change may be performed again.
[0203] In this way, there is no need to prepare replacement hardware resources for hardware resources that are not the cause of application instability, which reduces waste of replacement hardware resources when replacing an unstable application.
[0204] Furthermore, in this embodiment, when an application is determined to be unstable, the policy manager unit 90 may estimate the cause of the instability of the application. If it is estimated that the cause of the instability of the application is the application itself, a setting change may be executed to remove at least one replacement candidate resource from the group of operating resources and add a new resource to the group of operating resources.
[0205] In this embodiment, the monitoring function unit 72 may calculate the stability evaluation value of a cluster based on the stability evaluation value of an application running in that cluster.
[0206] The policy manager unit 90 may then determine whether the cluster is unstable based on the stability evaluation value of the cluster. If the cluster is determined to be unstable, the policy manager unit 90 may replace all applications running on the cluster with other clusters.
[0207] Here, an example of the flow of processing related to estimation of the cause of application instability, which is performed in the platform system 30 according to this embodiment, will be described with reference to the flow diagram shown in FIG.
[0208] In this processing example, for example, the policy manager unit 90 monitors whether stability evaluation value data indicating the stability of the application is output to the data bus unit 68 (S101).
[0209] When it is detected that the stability evaluation value data has been output to the data bus unit 68, the policy manager unit 90 acquires the stability evaluation value data (S102).
[0210] Then, the policy manager unit 90 determines whether the application is unstable or not based on the stability evaluation value data indicating the stability of the application, which is acquired in the process shown in S102 (S103).
[0211] If it is not determined to be unstable (S103: N), the process returns to S101.
[0212] If it is determined that the application is unstable (S103: Y), the policy manager unit 90 identifies the multiple hardware resources on which the processes included in the application are running (S104). In this processing example, the multiple hardware resources on which the processes included in the application are running can be identified by referring to the inventory data.
[0213] The policy manager unit 90 then identifies an application whose process is running on at least one of the hardware resources identified in the process shown in S104 (S105). In this process example, by referencing the inventory data, it is possible to identify, for each of the multiple hardware resources, an application whose process is running on that hardware resource.
[0214] Then, the policy manager unit 90 identifies the latest application stability evaluation value of at least one application identified in the process shown in S105 (S106).
[0215] Then, based on the application stability evaluation value identified in the process shown in S106, the cause of the application determined to be unstable in the process shown in S103 is estimated (S107).
[0216] The policy manager unit 90 then executes an action according to the cause estimated in the process of S107 (S108). In the process of S108, for example, the policy manager unit 90, the life cycle management unit 94, the container management unit 78, and the configuration management unit 76 may cooperate with one another to execute the action. Then, the process returns to the process of S101.
[0217] Next, an example of the processing flow for gradually replacing processes included in an application determined to be unstable, which is performed in the platform system 30 according to this embodiment, will be described with reference to the flow diagram illustrated in FIG. 13.
[0218] In this processing example, it is assumed that an application is determined to be unstable based on a stability evaluation value that indicates the stability of an application whose processes are distributed and running across a group of operating resources that includes multiple replacement candidate resources.
[0219] In this case, the policy manager unit 90 removes at least one replacement candidate resource from the group of operating resources and executes a setting change to add a new resource to the group of operating resources (S201).
[0220] Then, the policy manager unit 90 checks whether all replacement candidate resources have been excluded from the group of operating resources (S202).
[0221] If all replacement candidate resources have been excluded from the group of operating resources (S202: Y), the processing shown in this processing example is terminated.
[0222] If all replacement candidate resources have not been excluded from the group of operating resources (S202: N), the policy manager unit 90 determines whether the application is unstable based on a stability evaluation value that indicates the stability of the application whose processes are running distributed across the group of operating resources after the setting change is executed in the processing shown in S201 (S203).
[0223] If it is determined in the process shown in S203 that the application is unstable (S203: Y), the process shown in S201 is executed again.
[0224] If it is determined in the process shown in S203 that the application is stable (S203: N), the process shown in this process example is terminated.
[0225] The present invention is not limited to the above-described embodiment.
[0226] For example, the functional units according to this embodiment are not limited to those shown in FIG.
[0227] Furthermore, the functional unit according to this embodiment does not need to be a NF in 5G. For example, the functional unit according to this embodiment may be a network node in 4G, such as an eNodeB, a vDU, a vCU, a Packet Data Network Gateway (P-GW), a Serving Gateway (S-GW), a Mobility Management Entity (MME), or a Home Subscriber Server (HSS).
[0228] Furthermore, the scope of application of the present invention is not limited to applications included in the communication system 1. The present invention is also applicable to general applications other than those included in the communication system 1.
[0229] Furthermore, the functional units according to the present embodiment may be realized using hypervisor-type or host-type virtualization technology instead of container-type virtualization technology. Furthermore, the functional units according to the present embodiment do not need to be implemented by software, but may be implemented by hardware such as electronic circuits. Furthermore, the functional units according to the present embodiment may be implemented by a combination of electronic circuits and software.
[0230] The technology described in the present disclosure can also be expressed as follows: [1] A replacement system including: instability determination means for determining whether an application is unstable based on a stability evaluation value indicating the stability of an application whose processes are distributed across a group of operating resources including a plurality of replacement candidate resources; and configuration change means for executing a configuration change to remove at least one of the replacement candidate resources from the group of operating resources and add a new resource to the group of operating resources if the application is determined to be unstable based on the stability evaluation value indicating the stability of the application whose processes are distributed across the group of operating resources after the configuration change is executed, wherein the configuration change means executes the configuration change again if the application is determined to be unstable based on the stability evaluation value indicating the stability of the application whose processes are distributed across the group of operating resources after the configuration change is executed. [2] The replacement system described in [1], wherein the instability determination means determines whether the application is unstable based on the stability evaluation value indicating the stability of the application whose processes are distributed across the group of operating resources after the configuration change is executed each time the configuration change is executed. [3] The replacement system according to [1], characterized in that the setting change means repeats the execution of the setting change until the application is determined to be stable based on a stability evaluation value indicating the stability of the application whose processes are running in a distributed manner across the group of operating resources, or until all of the replacement candidate resources are excluded from the group of operating resources. [4] The replacement system according to any one of [1] to [3], characterized in that the instability determination means determines whether the application is unstable based on a stability evaluation value indicating the stability of the application whose processes are running in a distributed manner across the group of operating resources after the setting change, in response to the passage of a predetermined time since the setting change was executed.[5] The replacement system according to any one of [1] to [4], further comprising: an exclusion resource determination means for determining a replacement candidate resource to be excluded from the group of operating resources based on a stability evaluation value indicating the stability of other applications running on each of the plurality of replacement candidate resources included in the group of operating resources; and the setting change means excludes the determined replacement candidate resource from the group of operating resources if it is determined that an application, whose processes are running in a distributed manner across the group of operating resources, is unstable based on the stability evaluation value indicating the stability of the application. [6] The replacement system according to any one of [1] to [5], characterized in that the process is an execution unit of the application in a container-type virtualized application execution environment. [7] The replacement system according to any one of [1] to [6], characterized in that the application is an application included in a communication system. [8] The replacement system according to [7], characterized in that the application is a network function. [9] A replacement method comprising: determining whether an application is unstable based on a stability evaluation value indicating the stability of an application whose processes are distributed across a group of operating resources including a plurality of replacement candidate resources; if the application is determined to be unstable based on the stability evaluation value indicating the stability of the application whose processes are distributed across the group of operating resources, executing a configuration change to exclude at least one of the replacement candidate resources from the group of operating resources and add a new resource to the group of operating resources; and if the application is determined to be unstable based on the stability evaluation value indicating the stability of the application whose processes are distributed across the group of operating resources after the configuration change has been executed, executing the configuration change again.
Claims
1. An instability determination means for determining whether an application is unstable based on a stability evaluation value indicating the stability of an application whose processes are distributed and running on a group of operating resources including a plurality of replacement candidate resources; and a setting change means for executing a setting change to remove at least one of the replacement candidate resources from the group of operating resources and add a new resource to the group of operating resources when the application is determined to be unstable based on a stability evaluation value indicating the stability of the application, the process of which is distributed among the group of operating resources; the setting change means executes the setting change again when it is determined that the application is unstable based on a stability evaluation value indicating stability of the application, the process of which is distributed among the group of operating resources after the setting change is executed. Replacement system.
2. the instability determination means determines whether or not the application is unstable based on a stability evaluation value indicating stability of the application, the process of which is distributed among the group of operating resources after the setting change is executed, each time the setting change is executed. The replacement system according to claim 1.
3. the setting change means repeats execution of the setting change until the application is determined to be stable based on a stability evaluation value indicating the stability of the application, the process of which is distributed among the group of operating resources, or until all of the replacement candidate resources are removed from the group of operating resources. The replacement system according to claim 1.
4. the instability determination means, in response to a lapse of a predetermined time since the setting change was executed, determines whether or not the application is unstable based on a stability evaluation value indicating stability of the application, the process of which is running distributed among the group of operating resources after the setting change was executed; The replacement system according to claim 1.
5. The present invention further includes an excluded resource determination means for determining a replacement candidate resource to be excluded from the group of operating resources based on a stability evaluation value indicating the stability of other applications running on each of the plurality of replacement candidate resources included in the group of operating resources, the setting change means, when it is determined that the application is unstable based on a stability evaluation value indicating the stability of the application, the process of which is distributed among the group of operating resources, excludes the determined replacement candidate resource from the group of operating resources; The replacement system according to claim 1.
6. The process is an execution unit of the application in a container-type virtualized application execution environment. The replacement system according to claim 1.
7. The application is an application included in a communication system. The replacement system according to claim 1.
8. The application is a network function. The replacement system according to claim 7.
9. determining whether an application is unstable based on a stability evaluation value indicating stability of the application, the process of which is distributed and operated among a group of operating resources including a plurality of replacement candidate resources; When it is determined that the application is unstable based on a stability evaluation value indicating the stability of the application, the process of which is distributed among the group of operating resources, executing a setting change to remove at least one of the replacement candidate resources from the group of operating resources and to add a new resource to the group of operating resources; executing the setting change again when it is determined that the application is unstable based on a stability evaluation value indicating the stability of the application, the process of which is distributed among the group of operating resources after the setting change is executed; Replacement methods including: