Possible causes of application instability
The cause estimation system distinguishes between application and hardware instability in communication systems, improving stability evaluation by identifying separate stability values for applications and their resources.
Patent Information
- Application Number
- JP2024566967
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-12-26
- Publication Date
- 2026-02-16
- Estimated Expiration
- 2042-12-26
AI Technical Summary
Existing methods for evaluating application stability in communication systems fail to distinguish between application instability caused by the application itself or hardware resources, leading to inaccurate cause identification.
A cause estimation system that identifies stability evaluation values for both the application and the hardware resources it is running on, allowing for determination of whether instability is due to the application or the hardware, using a first and second stability evaluation value identification means.
Accurately estimates the cause of application instability, enabling targeted improvements in hardware or application performance.
Smart Images

Figure 0007814556000001 
Figure 0007814556000002 
Figure 0007814556000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to estimating the cause of application instability. [Background technology]
[0002] Patent Document 1 describes deploying a network function included in a communication system to a server on which a container-type application execution environment is installed. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] International Publication No. 2021 / 171210 Summary of the Invention [Problem to be solved by the invention]
[0004] Some applications, such as network functions, that operate in a communication system include many processes, and these processes may be distributed and run on multiple hardware resources.
[0005] When evaluating the stability of such an application, determining whether each of the many processes included in the application is unstable based on a stability evaluation value indicating the stability of the process would require an enormous amount of calculation.
[0006] Therefore, when evaluating the stability of such an application, it is common to determine whether the application is unstable based on a stability evaluation value aggregated for multiple processes. For example, for each type of process included in the application, it is determined whether the application is unstable based on a stability evaluation value aggregated for multiple processes of that type.
[0007] However, when an application is determined to be unstable based on a stability evaluation value aggregated across multiple processes, since the stability evaluation value is an aggregate value across multiple hardware resources, it is not possible to determine from the stability evaluation value alone whether the cause of the application's instability is the application itself or the hardware resources.
[0008] The above is not limited to applications included in communication systems, but also applies to general applications.
[0009] The present invention has been made in view of the above-mentioned circumstances, and one of its objects is to make it possible to accurately estimate the cause of application instability. [Means for solving the problem]
[0010] In order to solve the above problem, the cause estimation system of the present disclosure includes a first stability evaluation value identification means for identifying a stability evaluation value indicating the stability of a first application whose included processes are distributed across multiple hardware resources and running; a second stability evaluation value identification means for identifying a stability evaluation value indicating the stability of a second application whose at least one process is running on the hardware resource on which at least one process included in the first application is running; an instability determination means for determining whether the application is unstable based on the stability evaluation value indicating the stability of the application; and a cause estimation means for, when the first application is determined to be unstable based on the stability evaluation value indicating the stability of the first application, estimating whether the instability of the first application is caused by the first application or by the hardware resource on which the process included in the first application is running, based on the stability evaluation value indicating the stability of the second application.
[0011] In addition, a cause estimation method according to the present disclosure includes identifying a stability evaluation value indicating the stability of a first application, the processes of which are distributed across multiple hardware resources and running; identifying a stability evaluation value indicating the stability of a second application, the processes of which are running on the hardware resources on which at least one process included in the first application is running; determining whether the application is unstable based on the stability evaluation value indicating the stability of the application; and, if it is determined that the first application is unstable based on the stability evaluation value indicating the stability of the first application, estimating whether the instability of the first application is caused by the first application or by the hardware resources on which the processes included in the first application are running, based on the stability evaluation value indicating the stability of the second application. [Brief explanation of the drawings]
[0012] [Figure 1] 1 is a diagram illustrating an example of a communication system according to an embodiment of the present invention. [Figure 2] 1 is a diagram illustrating an example of a communication system according to an embodiment of the present invention. [Figure 3] FIG. 1 is a diagram illustrating an example of a network service according to an embodiment of the present invention. [Figure 4] FIG. 1 is a diagram illustrating an example of associations between elements established in a communication system according to an embodiment of the present invention. [Figure 5] FIG. 2 is a functional block diagram showing an example of functions implemented in a platform system according to an embodiment of the present invention. [Figure 6] FIG. 2 illustrates an example of a data structure of physical inventory data. [Figure 7] FIG. 1 is a diagram illustrating an example of a situation in which processes included in a plurality of applications are distributed and run on a plurality of hardware resources. [Figure 8]FIG. 1 is a diagram illustrating an example of a situation in which processes included in a plurality of applications are distributed and run on a plurality of hardware resources. [Figure 9] FIG. 2 is a flowchart showing an example of a flow of processing performed in a platform system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0013] Hereinafter, an embodiment of the present invention will be described in detail with reference to the drawings.
[0014] 1 and 2 are diagrams illustrating an example of a communication system 1 according to an embodiment of the present invention. Fig. 1 is a diagram focusing on the locations of a group of data centers included in the communication system 1. Fig. 2 is a diagram focusing on various computer systems implemented in the group of data centers included in the communication system 1.
[0015] As shown in FIG. 1, the data centers included in the communication system 1 are classified into a central data center 10, regional data centers 12, and edge data centers .
[0016] For example, several central data centers 10 are distributed and placed within the area covered by the communication system 1 (for example, within Japan).
[0017] For example, several tens of regional data centers 12 are distributed and placed within the area covered by the communication system 1. For example, if the area covered by the communication system 1 is the entire country of Japan, one or two regional data centers 12 may be placed in each prefecture.
[0018] For example, several thousand edge data centers 14 are distributed within the area covered by the communication system 1. Each edge data center 14 is capable of communicating with communication equipment 18 equipped with an antenna 16. As shown in FIG. 1, one edge data center 14 may be capable of communicating with several communication equipment 18. The communication equipment 18 may include a computer such as a server computer. The communication equipment 18 according to this embodiment performs wireless communication with a UE (User Equipment) 20 via the antenna 16. The communication equipment 18 equipped with the antenna 16 is provided with, for example, a radio unit (RU) (described later).
[0019] In the central data center 10, the regional data center 12, and the edge data center 14 according to this embodiment, multiple servers are arranged.
[0020] In this embodiment, for example, the central data center 10, the regional data centers 12, and the edge data centers 14 are capable of communicating with each other. Furthermore, the central data centers 10, the regional data centers 12, and the edge data centers 14 are also capable of communicating with each other.
[0021] 2, the communication system 1 according to this embodiment includes a platform system 30, multiple radio access networks (RANs) 32, multiple core network systems 34, and multiple UEs 20. The core network systems 34, the RANs 32, and the UEs 20 cooperate with each other to realize a mobile communication network.
[0022] The RAN 32 is a computer system equipped with an antenna 16, which corresponds to an eNodeB (eNB) in a fourth-generation mobile communication system (hereinafter referred to as 4G) or a gNB (NR base station) in a fifth-generation mobile communication system (hereinafter referred to as 5G). The RAN 32 according to this embodiment is mainly implemented by a group of servers and communication equipment 18 arranged in an edge data center 14. Note that part of the RAN 32 (for example, a distributed unit (DU), a central unit (CU), a virtual distributed unit (vDU), and a virtual central unit (vCU)) may be implemented in the central data center 10 or the regional data center 12, rather than in the edge data center 14.
[0023] The core network system 34 is a system equivalent to an EPC (Evolved Packet Core) in 4G or a 5G Core (5GC) in 5G. The core network system 34 according to this embodiment is implemented mainly by a group of servers arranged in the central data center 10 and the regional data centers 12.
[0024] The platform system 30 according to this embodiment is configured on, for example, a cloud platform, and includes a processor 30a, a storage unit 30b, and a communication unit 30c, as shown in FIG. 2. The processor 30a is a program-controlled device such as a microprocessor that operates according to a program installed in the platform system 30. The storage unit 30b is, for example, a storage element such as a ROM or RAM, a solid-state drive (SSD), or a hard disk drive (HDD). The storage unit 30b stores programs executed by the processor 30a. The communication unit 30c is, for example, a communication interface such as a network interface controller (NIC) or a wireless local area network (LAN) module. Note that software-defined networking (SDN) may be implemented in the communication unit 30c. The communication unit 30c exchanges data with the RAN 32 and the core network system 34.
[0025] In this embodiment, the platform system 30 is implemented by a group of servers located in the central data center 10. Note that the platform system 30 may also be implemented by a group of servers located in the regional data centers 12.
[0026] In this embodiment, for example, in response to a purchase request for a network service (NS) from a purchaser, the requested network service is established in the RAN 32 and the core network system 34. Then, the established network service is provided to the purchaser.
[0027] For example, a purchaser such as an MVNO (Mobile Virtual Network Operator) is provided with network services such as voice communication services and data communication services. The voice communication services and data communication services provided by this embodiment are ultimately provided to customers (end users) of the purchaser (MVNO in the above example) who use the UE 20 shown in FIGS. 1 and 2. The end users can perform voice communication and data communication with other users via the RAN 32 and the core network system 34. The UE 20 of the end user can also access a data network such as the Internet via the RAN 32 and the core network system 34.
[0028] In addition, in this embodiment, an IoT (Internet of Things) service may be provided to an end user who uses a robot arm, a connected car, etc. In this case, for example, the end user who uses the robot arm, the connected car, etc. may become a purchaser of the network service according to this embodiment.
[0029] In this embodiment, a container-based virtualized application execution environment such as Docker (registered trademark) is installed on servers located in the central data center 10, the regional data centers 12, and the edge data center 14, allowing containers to be deployed and run on these servers. A cluster consisting of one or more containers generated by such virtualization technology may be built on these servers. For example, a Kubernetes cluster managed by a container management tool such as Kubernetes (registered trademark) may be built. Then, a processor on the built cluster may execute a container-based application.
[0030] In this embodiment, the network service provided to the purchaser is composed of one or more functional units (for example, network functions (NFs)). In this embodiment, the functional units are implemented as NFs realized by virtualization technology. NFs realized by virtualization technology are called VNFs (Virtualized Network Functions). It does not matter what virtualization technology is used to virtualize them. For example, in this description, a CNF (Containerized Network Function) realized by container-type virtualization technology is also included in the VNF. In this embodiment, the network service will be described as being implemented by one or more CNFs. Furthermore, the functional units in this embodiment may correspond to network nodes.
[0031] Fig. 3 is a diagram illustrating an example of an operating network service. The network service illustrated in Fig. 3 includes, as software elements, NFs such as a plurality of RUs 40, a plurality of DUs 42, a plurality of CUs 44 (CU-CP (Central Unit - Control Plane) 44a and CU-UP (Central Unit - User Plane) 44b), a plurality of AMFs (Access and Mobility Management Functions) 46, a plurality of SMFs (Session Management Functions) 48, and a plurality of UPFs (User Plane Functions) 50.
[0032] In the example of Figure 3, RU 40, DU 42, CU-CP 44a, AMF 46, and SMF 48 correspond to elements of the control plane (C-Plane), and RU 40, DU 42, CU-UP 44b, and UPF 50 correspond to elements of the user plane (U-Plane).
[0033] The network service may include other types of NFs as software elements. The network service is implemented on computer resources (hardware elements) such as multiple servers.
[0034] In this embodiment, for example, a communication service in a certain area is provided by the network service shown in FIG.
[0035] In this embodiment, it is assumed that multiple RUs 40, multiple DUs 42, multiple CU-UPs 44b, and multiple UPFs 50 shown in FIG. 3 belong to one end-to-end network slice.
[0036] Fig. 4 is a diagram schematically illustrating an example of associations between elements established in the communication system 1 in this embodiment. The symbols M and N shown in Fig. 4 represent any integers equal to or greater than 1, and indicate the relationship between the numbers of elements connected by a link. When both ends of a link are a combination of M and N, the elements connected by the link have a many-to-many relationship, and when both ends of a link are a combination of 1 and N or a combination of 1 and M, the elements connected by the link have a one-to-many relationship.
[0037] As shown in Figure 4, network services (NS), network functions (NF), CNFCs (Containerized Network Function Components), pods, and containers have a hierarchical structure.
[0038] An NS corresponds to, for example, a network service configured from multiple NFs. Here, an NS may correspond to an element of granularity such as a 5G RAN (gNB), an EPC, a 5G RAN (eNB), or the like.
[0039] In 5G, NFs correspond to elements with granularity such as RU, DU, CU-CP, CU-UP, AMF, SMF, and UPF. In 4G, NFs correspond to elements with granularity such as MME (Mobility Management Entity), HSS (Home Subscriber Server), S-GW (Serving Gateway), vDU, and vCU. In this embodiment, for example, one NS includes one or more NFs. In other words, one or more NFs are subordinate to one NS.
[0040] A CNFC corresponds to an element of granularity such as DU mgmt or DU Processing. A CNFC may be a microservice deployed on a server as one or more containers. For example, a CNFC may be a microservice that provides some of the functions of DU, CU-CP, CU-UP, etc. Also, a CNFC may be a microservice that provides some of the functions of UPF, AMF, SMF, etc. In this embodiment, for example, one NF includes one or more CNFCs. In other words, one or more CNFCs are subordinate to one NF.
[0041] A pod refers to the smallest unit for managing a Docker container in Kubernetes, for example. In this embodiment, for example, one CNFC includes one or more pods. In other words, one or more pods are under the control of one CNFC.
[0042] In this embodiment, for example, one pod includes one or more containers. That is, one or more containers are subordinate to one pod.
[0043] Also, as shown in Figure 4, network slices (NSIs) and network slice subnet instances (NSSIs) have a hierarchical structure.
[0044] The NSI can also be considered an end-to-end virtual circuit spanning multiple domains (e.g., from the RAN 32 to the core network system 34). The NSI may be a slice for high-speed, high-capacity communication (e.g., for enhanced Mobile Broadband (eMBB)), a slice for high-reliability and low-latency communication (e.g., for Ultra-Reliable and Low Latency Communications (URLLC)), or a slice for connecting a large number of terminals (e.g., for massive Machine Type Communication (mMTC)). The NSSI can also be considered a virtual circuit of a single domain obtained by dividing the NSI. The NSSI may be a slice of the RAN domain, a slice of a transport domain such as the Mobile Back Haul (MBH) domain, or a slice of the core network domain.
[0045] In this embodiment, for example, one NSI includes one or more NSSIs. That is, one or more NSSIs are subordinate to one NSI. Note that in this embodiment, multiple NSIs may share the same NSSI.
[0046] Furthermore, as shown in FIG. 4, NSSI and NS generally have a many-to-many relationship.
[0047] Furthermore, in this embodiment, for example, one NF can belong to one or more network slices. Specifically, for example, one NF can be configured with NSSAI (Network Slice Selection Assistance Information) including one or more S-NSSAI (Sub Network Slice Selection Assist Information). Here, S-NSSAI is information associated with a network slice. Note that an NF does not necessarily have to belong to a network slice.
[0048] Fig. 5 is a functional block diagram showing an example of functions implemented in the platform system 30 according to this embodiment. Note that the platform system 30 according to this embodiment does not need to implement all of the functions shown in Fig. 5, and functions other than the functions shown in Fig. 5 may also be implemented.
[0049] As shown in FIG. 5 , the platform system 30 according to this embodiment functionally includes, for example, an operations support system (OSS) unit 60, an orchestration (E2EO: End-to-End-Orchestration) unit 62, a service catalog storage unit 64, a big data platform unit 66, a data bus unit 68, an AI (Artificial Intelligence) unit 70, a monitoring function unit 72, an SDN controller 74, a configuration management unit 76, a container management unit 78, and a repository unit 80. The OSS unit 60 includes an inventory database 82, a ticket management unit 84, a fault management unit 86, and a performance management unit 88. The E2EO unit 62 includes a policy manager unit 90, a slice manager unit 92, and a lifecycle management unit 94. These elements are implemented primarily using a processor 30 a, a storage unit 30 b, and a communication unit 30 c.
[0050] The functions shown in Fig. 5 may be implemented by having a processor 30a execute a program that is installed in a platform system 30, which is one or more computers, and that includes instructions corresponding to the functions. This program may be supplied to the platform system 30 via a computer-readable information storage medium, such as an optical disk, a magnetic disk, a magnetic tape, a magneto-optical disk, or a flash memory, or via the Internet. The functions shown in Fig. 5 may also be implemented using circuit blocks, memory, or other LSIs. Those skilled in the art will understand that the functions shown in Fig. 5 can be realized in various forms, such as hardware alone, software alone, or a combination thereof.
[0051] The container management unit 78 manages the lifecycle of a container, which includes, for example, processes related to the construction of a container, such as the deployment and configuration of the container.
[0052] Here, the platform system 30 according to this embodiment may include a plurality of container management units 78. A container management tool such as Kubernetes and a package manager such as Helm may be installed in each of the plurality of container management units 78. Each of the plurality of container management units 78 may execute container construction, such as container deployment, for a server group (e.g., a Kubernetes cluster) associated with the corresponding container management unit 78.
[0053] It should be noted that the container management unit 78 does not need to be included in the platform system 30. The container management unit 78 may be provided, for example, in a server managed by the container management unit 78 (i.e., the RAN 32 or the core network system 34), or may be provided in another server that is annexed to the server managed by the container management unit 78.
[0054] In this embodiment, the repository unit 80 stores, for example, container images of containers included in a group of functional units (for example, a group of NFs) that realize a network service.
[0055] The inventory database 82 is a database that stores inventory information, which includes, for example, information about servers that are installed in the RAN 32 and the core network system 34 and that are managed by the platform system 30.
[0056] In this embodiment, inventory data is stored in the inventory database 82. The inventory data indicates the configuration of the elements included in the communication system 1 and the current status of the associations between the elements. The inventory data also indicates the status of resources managed by the platform system 30 (for example, the usage status of the resources). The inventory data may be physical inventory data or logical inventory data. The physical inventory data and logical inventory data will be described later.
[0057] Fig. 6 is a diagram showing an example of the data structure of physical inventory data. The physical inventory data shown in Fig. 6 is associated with one server. The physical inventory data shown in Fig. 6 includes, for example, a server ID, location data, building data, floor data, rack data, specification data, network data, an operating container ID list, a cluster ID, and the like.
[0058] The server ID included in the physical inventory data is, for example, an identifier of the server associated with the physical inventory data.
[0059] The location data included in the physical inventory data is, for example, data indicating the location (for example, the address of the location) of the server associated with the physical inventory data.
[0060] The building data included in the physical inventory data is, for example, data indicating the building (for example, the building name) in which the server associated with the physical inventory data is located.
[0061] The floor number data included in the physical inventory data is, for example, data indicating the floor number on which the server associated with the physical inventory data is located.
[0062] The rack data included in the physical inventory data is, for example, an identifier of the rack in which the server associated with the physical inventory data is located.
[0063] The specification data included in the physical inventory data is, for example, data indicating the specifications of the server associated with the physical inventory data, and the specification data indicates, for example, the number of cores, memory capacity, hard disk capacity, etc.
[0064] The network data included in the physical inventory data is, for example, data indicating information about the network of the server associated with the physical inventory data, and the network data indicates, for example, the NICs that the server has, the number of ports that the NICs have, the port IDs of the ports, etc.
[0065] The operating container ID list included in the physical inventory data is, for example, data that indicates information about one or more containers operating on a server associated with the physical inventory data, and the operating container ID list indicates, for example, a list of identifiers (container IDs) of instances of the containers.
[0066] The cluster ID included in the physical inventory data is, for example, an identifier of a cluster (for example, a Kubernetes cluster) to which a server associated with the physical inventory data belongs.
[0067] The logical inventory data includes topology data indicating the current state of associations between multiple elements included in the communication system 1, such as those shown in Fig. 4. For example, the logical inventory data includes topology data including an identifier of a certain NS and identifiers of one or more NFs under the NS. Also, for example, the logical inventory data includes topology data including an identifier of a certain network slice and identifiers of one or more NFs belonging to the network slice.
[0068] The inventory data may also include data indicating the current status of the geographical relationships and topological relationships between the elements included in the communication system 1. As described above, the inventory data includes location data indicating the locations where the elements included in the communication system 1 are operating, i.e., the current locations of the elements included in the communication system 1. From this, it can be said that the inventory data indicates the current status of the geographical relationships between the elements (for example, the geographical proximity between the elements).
[0069] The logical inventory data may also include NSI data indicating information about the network slice. The NSI data indicates attributes such as an identifier of an instance of the network slice and a type of the network slice. The logical inventory data may also include NSSI data indicating information about the network slice subnet. The NSSI data indicates attributes such as an identifier of an instance of the network slice subnet and a type of the network slice subnet.
[0070] The logical inventory data may also include NS data indicating information about an NS. The NS data indicates attributes such as an NS instance identifier and an NS type. The logical inventory data may also include NF data indicating information about an NF. The NF data indicates attributes such as an NF instance identifier and an NF type. The logical inventory data may also include CNFC data indicating information about a CNFC. The CNFC data indicates attributes such as an instance identifier and a CNFC type. The logical inventory data may also include pod data indicating information about a pod included in the CNFC. The pod data indicates attributes such as a pod instance identifier and a pod type. The logical inventory data may also include container data indicating information about a container included in the pod. The container data indicates attributes such as a container ID of a container instance and a container type.
[0071] The container ID of the container data included in the logical inventory data and the container ID included in the operating container ID list included in the physical inventory data associate a container instance with the server on which the container instance is running.
[0072] Furthermore, the logical inventory data may include data indicating various attributes such as a host name and an IP address. For example, container data may include data indicating an IP address of a container corresponding to the container data. For example, NF data may include data indicating an IP address and a host name of the NF indicated by the NF data.
[0073] The logical inventory data may also include data indicating an NSSAI, including one or more S-NSSAIs, configured in each NF.
[0074] Furthermore, the inventory database 82 is able to grasp the resource status as needed in cooperation with the container management unit 78. The inventory database 82 then updates the inventory data stored therein as needed based on the latest resource status.
[0075] In addition, in response to actions being performed, such as constructing a new element included in the communication system 1, changing the configuration of an element included in the communication system 1, scaling an element included in the communication system 1, or replacing an element included in the communication system 1, the inventory database 82 updates the inventory data stored in the inventory database 82.
[0076] The service catalog storage unit 64 stores service catalog data. The service catalog data may include, for example, service template data indicating logic used by the life cycle management unit 94. This service template data includes information necessary for building a network service. For example, the service template data includes information defining NSs, NFs, and CNFCs, and information indicating the correspondence between NSs, NFs, and CNFCs. Furthermore, for example, the service template data includes a workflow script for building a network service.
[0077] An example of service template data is an NSD (NS Descriptor). The NSD is associated with a network service and indicates the types of multiple functional units (e.g., multiple CNFs) included in the network service. The NSD may also indicate the number of each type of functional unit, such as CNF, included in the network service. The NSD may also indicate the file names of CNFDs (described later) related to the CNFs included in the network service.
[0078] An example of service template data is a CNF Descriptor (CNFD). The CNFD may indicate the computer resources (e.g., CPU, memory, hard disk, etc.) required by the CNF. For example, the CNFD may indicate the computer resources (CPU, memory, hard disk, etc.) required by each of multiple containers included in the CNF.
[0079] The service catalog data may also include information about thresholds (for example, anomaly detection thresholds) that are used by the policy manager 90 to compare with the calculated performance index values. The performance index values will be described later.
[0080] The service catalog data may also include, for example, slice template data, which includes information necessary to perform instantiation of a network slice, including, for example, logic utilized by the slice manager unit 92.
[0081] The slice template data includes information on the "Generic Network Slice Template" defined by the GSMA (GSM Association) ("GSM" is a registered trademark). Specifically, the slice template data includes network slice template data (NST), network slice subnet template data (NSST), and network service template data. The slice template data also includes information indicating the hierarchical structure of these elements, as shown in FIG. 4.
[0082] In this embodiment, for example, the life cycle management unit 94 constructs a new network service in response to a purchase request for an NS from a purchaser.
[0083] For example, in response to a purchase request, the lifecycle management unit 94 may execute a workflow script associated with the network service to be purchased. By executing this workflow script, the lifecycle management unit 94 may instruct the container management unit 78 to deploy a container included in the new network service to be purchased. The container management unit 78 may then obtain a container image of the container from the repository unit 80 and deploy the container corresponding to the container image to a server.
[0084] In addition, in this embodiment, the life cycle management unit 94 executes, for example, scaling and replacement of elements included in the communication system 1. Here, the life cycle management unit 94 may output a container deployment instruction or deletion instruction to the container management unit 78. Then, the container management unit 78 may execute processing such as container deployment or container deletion in accordance with the instruction. In this embodiment, the life cycle management unit 94 is capable of executing scaling and replacement that cannot be handled by a tool such as Kubernetes in the container management unit 78.
[0085] Furthermore, the life cycle management unit 94 may output an instruction to create a communication path to the SDN controller 74. For example, the life cycle management unit 94 presents two IP addresses at both ends of the communication path to be created to the SDN controller 74, and the SDN controller 74 creates a communication path connecting these two IP addresses. The created communication path may be managed in association with these two IP addresses.
[0086] Furthermore, the life cycle management unit 94 may output to the SDN controller 74 an instruction to create a communication path between the two IP addresses that is associated with the two IP addresses.
[0087] In this embodiment, for example, the slice manager unit 92 performs instantiation of a network slice. In this embodiment, for example, the slice manager unit 92 performs instantiation of a network slice by executing logic indicated by a slice template stored in the service catalog storage unit 64.
[0088] The slice manager unit 92 is configured to include the functions of the NSMF (Network Slice Management Function) and the NSSMF (Network Slice Sub-network Management Function) described in, for example, the 3GPP (registered trademark) (Third Generation Partnership Project) specification "TS28 533." The NSMF is a function that generates and manages network slices and provides NSI management services. The NSSMF is a function that generates and manages network slice subnets that constitute part of the network slice and provides NSSI management services.
[0089] Here, the slice manager unit 92 may output a configuration management instruction related to instantiation of the network slice to the configuration management unit 76. Then, the configuration management unit 76 may perform configuration management such as setting in accordance with the configuration management instruction.
[0090] Furthermore, the slice manager unit 92 may present two IP addresses to the SDN controller 74 and output an instruction to create a communication path between these two IP addresses.
[0091] In this embodiment, the configuration management unit 76 executes configuration management such as setting of element groups such as NFs in accordance with configuration management instructions received from the life cycle management unit 94 and the slice manager unit 92, for example.
[0092] In this embodiment, the SDN controller 74 creates a communication path between two IP addresses associated with a communication path creation instruction received from, for example, the life cycle management unit 94 or the slice manager unit 92. The SDN controller 74 may create a communication path between two IP addresses using a known path calculation method such as Flex Algo.
[0093] For example, the SDN controller 74 may use a segment routing technology (e.g., SRv6 (segment routing IPv6)) to construct NSIs and NSSIs for aggregation routers, servers, and the like present along the communication paths. The SDN controller 74 may also generate NSIs and NSSIs across multiple target NFs by issuing commands to multiple target NFs to set up a common Virtual Local Area Network (VLAN) and commands to assign the bandwidth and priority indicated in the setting information to the VLAN.
[0094] In addition, the SDN controller 74 may perform operations such as changing the maximum bandwidth available for communication between two IP addresses without constructing a network slice.
[0095] The platform system 30 according to this embodiment may include multiple SDN controllers 74. Each of the multiple SDN controllers 74 may execute processing such as creating a communication path for a group of network devices such as an AG associated with the SDN controller 74.
[0096] In this embodiment, for example, the monitoring function unit 72 monitors the group of elements included in the communication system 1 in accordance with a given management policy. Here, the monitoring function unit 72 may monitor the group of elements in accordance with a monitoring policy specified by a purchaser when purchasing a network service, for example.
[0097] In this embodiment, the monitoring function unit 72 performs monitoring at various levels, such as the slice level, the NS level, the NF level, the CNFC level, and the hardware level of a server or the like.
[0098] For example, the monitoring function unit 72 may set a module that outputs metric data in hardware such as a server or in a software element included in the communication system 1 so that monitoring can be performed at the various levels described above. Here, for example, an NF may output metric data indicating metrics that are measurable (identifiable) in the NF to the monitoring function unit 72. Also, a server may output metric data indicating metrics related to hardware that is measurable (identifiable) in the server to the monitoring function unit 72.
[0099] Furthermore, for example, the monitoring function unit 72 may deploy a sidecar container on the server that aggregates metric data indicating metrics output from multiple containers for each CNFC (microservice). This sidecar container may include an agent called an exporter. The monitoring function unit 72 may repeatedly execute, at a given monitoring interval, a process of acquiring metric data aggregated for each microservice from the sidecar container, using the mechanisms of a monitoring tool such as Prometheus that can monitor container management tools such as Kubernetes.
[0100] The monitoring function unit 72 may monitor performance indicator values for performance indicators described in, for example, “TS 28.552, Management and orchestration; 5G performance measurements” or “TS 28.554, Management and orchestration; 5G end to end Key Performance Indicators (KPI).” Then, the monitoring function unit 72 may obtain metric data indicating the monitored performance indicator values.
[0101] In this embodiment, the monitoring function unit 72 performs a process (enrichment) to aggregate metric data, for example, in a predetermined aggregation unit, thereby generating performance index value data indicating the performance index values of the elements included in the communication system 1 in that aggregation unit.
[0102] For example, for one gNB, performance index value data for the gNB is generated by aggregating metric data indicating metrics of elements (e.g., network nodes such as DU42 and CU44) under the control of the gNB. In this way, performance index value data indicating communication performance in the area covered by the gNB is generated. Here, for example, performance index value data indicating multiple types of communication performance such as traffic volume (throughput) and latency may be generated for each gNB. Note that the communication performance indicated by the performance index value data is not limited to traffic volume and latency.
[0103] Then, the monitoring function unit 72 outputs the performance index value data generated by the above-mentioned enrichment to the data bus unit 68.
[0104] In this embodiment, for example, the data bus unit 68 receives performance index value data output from the monitoring function unit 72. Then, based on the received one or more pieces of performance index value data, the data bus unit 68 generates a performance index value file including the one or more pieces of performance index value data. Then, the data bus unit 68 outputs the generated performance index value file to the big data platform unit 66.
[0105] In this embodiment, the monitoring function unit 72 identifies a stability evaluation value indicating the stability of each application by, for example, executing a process (enrichment) of aggregating metric data related to the application. Then, the monitoring function unit 72 generates stability evaluation value data indicating the identified stability evaluation value.
[0106] Then, the monitoring function unit 72 outputs the generated stability evaluation value data to the data bus unit 68.
[0107] In this embodiment, the data bus unit 68 receives, for example, stability evaluation value data output from the monitoring function unit 72 .
[0108] In addition, elements such as network slices, NSs, NFs, CNFCs, and hardware such as servers included in the communication system 1 notify the monitoring function unit 72 of various alerts (for example, notification of an alert triggered by the occurrence of a failure).
[0109] Then, for example, when the monitoring function unit 72 receives the above-mentioned alert notification, it outputs alert message data indicating the notification to the data bus unit 68. Then, the data bus unit 68 generates an alert file in which alert message data indicating one or more notifications are compiled into a single file, and outputs the alert file to the big data platform unit 66.
[0110] In this embodiment, the big data platform unit 66 accumulates, for example, performance index value files and alert files output from the data bus unit 68.
[0111] In this embodiment, for example, a plurality of trained machine learning models are stored in advance in the AI unit 70. The AI unit 70 uses the various machine learning models stored in the AI unit 70 to perform estimation processing such as future prediction processing of the usage status and service quality of the communication system 1. The AI unit 70 may generate estimation result data indicating the results of the estimation processing.
[0112] The AI unit 70 may perform estimation processing based on the files stored in the big data platform unit 66 and the above-mentioned machine learning model. This estimation processing is suitable for infrequently predicting long-term trends.
[0113] The AI unit 70 is also capable of acquiring performance index value data stored in the data bus unit 68. The AI unit 70 may perform estimation processing based on the performance index value data stored in the data bus unit 68 and the above-described machine learning model. This estimation processing is suitable for performing short-term predictions frequently.
[0114] In this embodiment, for example, the performance management unit 88 calculates a performance index value (e.g., KPI) based on a plurality of metric data and the metrics indicated by the metric data. The performance management unit 88 may also calculate a performance index value that is an overall evaluation of a plurality of types of metrics (e.g., a performance index value related to an end-to-end network slice) that cannot be calculated from a single metric data. The performance management unit 88 may also generate overall performance index value data that indicates the performance index value that is the overall evaluation.
[0115] The performance management unit 88 may acquire the above-mentioned performance index value file from the big data platform unit 66. The performance management unit 88 may also acquire estimation result data from the AI unit 70. Then, performance index values such as KPIs may be calculated based on at least one of the performance index value file and the estimation result data. The performance management unit 88 may also directly acquire metric data from the monitoring function unit 72. Then, performance index values such as KPIs may be calculated based on the metric data.
[0116] In this embodiment, the fault management unit 86 detects the occurrence of a fault in the communication system 1 based on, for example, at least one of the above-mentioned metric data, the above-mentioned alert notification, the above-mentioned estimation result data, and the above-mentioned overall performance index value data. The fault management unit 86 may detect the occurrence of a fault that cannot be detected from a single piece of metric data or a single alert notification, for example, based on a predetermined logic. The fault management unit 86 may generate detected fault data that indicates the detected fault.
[0117] The fault management unit 86 may obtain metric data and alert notifications directly from the monitoring function unit 72. The fault management unit 86 may also obtain performance index value files and alert files from the big data platform unit 66. The fault management unit 86 may also obtain alert message data from the data bus unit 68.
[0118] In this embodiment, the policy manager unit 90 executes a predetermined judgment process based on at least one of the above-mentioned metric data, the above-mentioned performance index value data, the above-mentioned stability evaluation value data, the above-mentioned alert message data, the above-mentioned performance index value file, the above-mentioned alert file, the above-mentioned estimation result data, the above-mentioned overall performance index value data, and the above-mentioned detected fault data.
[0119] The policy manager unit 90 may then execute an action according to the result of the determination process. For example, the policy manager unit 90 may output a command to construct a network slice to the slice manager unit 92. The policy manager unit 90 may also output a command to scale or replace an element to the life cycle management unit 94 according to the result of the determination process.
[0120] The policy manager unit 90 according to this embodiment is capable of acquiring performance index value data stored in the data bus unit 68. The policy manager unit 90 may then execute a predetermined determination process based on the performance index value data acquired from the data bus unit 68. The policy manager unit 90 may also execute a predetermined determination process based on alert message data stored in the data bus unit 68.
[0121] Furthermore, the policy manager unit 90 according to this embodiment is capable of acquiring stability evaluation value data stored in the data bus unit 68. The policy manager unit 90 may then execute a predetermined determination process based on the stability evaluation value data acquired from the data bus unit 68. For example, the policy manager unit 90 may determine whether an application is unstable based on the stability evaluation value data indicating the stability of the application.
[0122] In this embodiment, for example, the ticket management unit 84 generates a ticket indicating the content to be notified to the administrator of the communication system 1. The ticket management unit 84 may generate a ticket indicating the content of the occurred fault data. The ticket management unit 84 may also generate a ticket indicating the values of performance index value data, stability evaluation value data, or metric data. The ticket management unit 84 may also generate a ticket indicating the determination result by the policy manager unit 90.
[0123] Then, the ticket management unit 84 notifies the administrator of the communication system 1 of the generated ticket. For example, the ticket management unit 84 may send an email with the generated ticket attached to the email address of the administrator of the communication system 1.
[0124] As described above, in this embodiment, the policy manager unit 90 determines whether an application is unstable based on a stability evaluation value indicating the stability of the application. If the application is determined to be unstable, the policy manager unit 90 estimates the cause of the application's instability. For example, the policy manager unit 90 estimates whether the cause of the application's instability lies in the application itself or in the hardware resources on which a process included in the application is running.
[0125] The process may be, for example, an execution unit (for example, a pod) of the application in a container-type virtualized application execution environment.
[0126] The application may also be a network function (e.g., DU42, CU-CP44a, CU-UP44b, AMF46, SMF48, UPF50, etc.).
[0127] The process of estimating the cause of application instability will be further described below.
[0128] In this embodiment, for example, as described above, the monitoring function unit 72 calculates a stability index value indicating the stability of each of a plurality of applications included in the communication system 1. These applications include a plurality of types of processes. For each type, a plurality of processes of that type run, thereby operating the entire application. Furthermore, for each type, the processes of that type run in a distributed manner across a plurality of hardware resources.
[0129] FIG. 7 is a diagram illustrating an example of a situation in which processes included in a plurality of applications are distributed and run on a plurality of hardware resources.
[0130] The example in FIG. 7 shows a situation in which four applications with identifiers AP1, AP2, AP3, and AP4 are running.
[0131] In this embodiment, for each type of application, a hardware resource on which the application of that type can run is predetermined. In the following description, the hardware resource is assumed to be a server, but the hardware resource does not have to be a server and may be, for example, a node.
[0132] Hereinafter, a hardware resource on which a certain type of application can run will be referred to as a tenant corresponding to that application.
[0133] 7 shows four servers with identifiers S1, S2, S3, and S4, respectively. These four servers belong to one cluster (for example, a Kubernetes cluster).
[0134] 7 are each assumed to be of a different type. The tenant corresponding to the application with the identifier AP1 includes servers with identifiers S1, S2, and S3. The tenant corresponding to the application with the identifier AP2 includes servers with identifiers S3 and S4. The tenant corresponding to the application with the identifier AP3 includes servers with identifiers S1, S2, S3, and S4. The tenant corresponding to the application with the identifier AP4 includes servers with identifiers S1 and S4.
[0135] 7 corresponds to one process (e.g., a pod). The numbers shown in the rounded rectangles are identifiers associated with the process types. That is, rounded rectangles with the same numbers correspond to processes of the same type.
[0136] As shown in Figure 7, an application with an identifier AP1 includes three types of processes with identifiers 1, 2, and 3. The three processes of the type with identifier 1 are running on servers with identifiers S1, S2, and S3, respectively. Three processes of the type with identifier 2 are running on servers with identifiers S1, S2, and S3, respectively. Two processes of the type with identifier 3 are running on servers with identifiers S1 and S2, respectively.
[0137] Furthermore, an application with an identifier AP2 includes four types of processes with identifiers 4, 5, 6, and 7. Two processes of the type with identifier 4 are running on servers with identifiers S3 and S4, respectively. One process of the type with identifier 5 is running on a server with identifier S3. Two processes of the type with identifier 6 are running on servers with identifiers S3 and S4, respectively. Two processes of the type with identifier 7 are running on servers with identifiers S3 and S4, respectively.
[0138] Furthermore, an application with an identifier AP3 includes three types of processes with identifiers 8, 9, and 10. Four processes of the type with identifier 8 are running on servers with identifiers S1, S2, S3, and S4, respectively. Three processes of the type with identifier 9 are running on servers with identifiers S1, S3, and S4, respectively. Three processes of the type with identifier 10 are running on servers with identifiers S2, S3, and S4, respectively.
[0139] Furthermore, an application with an identifier AP4 includes three types of processes with identifiers 11, 12, and 13. Two processes of the type with identifier 11 are running on servers with identifiers S1 and S4, respectively. Two processes of the type with identifier 12 are running on servers with identifiers S1 and S4, respectively. One process of the type with identifier 13 is running on the server with identifier S4.
[0140] In this embodiment, for example, the container management unit 78 controls each type of process so that it runs in a distributed manner across as many hardware resources as possible.
[0141] In this embodiment, for example, the monitoring function unit 72 acquires a value (metric) indicating the stability of each type of process. For example, the value indicates the state of the process (e.g., kube_pod_status_ready), the start time of the process (e.g., kube_pod_start_time), the length of time the process performed input / output (e.g., Metrics such as the number of dropped packets for a process (e.g., container_network_transmit_packets_dropped_total), the number of dropped packets for a process (e.g., container_network_receive_packets_dropped_total), and the number of dropped packets for a process (e.g., container_network_receive_packets_dropped_total) may be obtained.
[0142] The monitoring function unit 72 then calculates a weighted sum of the acquired metrics using weights associated with the respective types as a stability evaluation value indicating the stability of the process of that type. Here, for example, weights for each type of metric may be predetermined for each process type. Then, a weighted sum of the acquired metrics using the predetermined weights may be calculated as a stability evaluation value indicating the stability of the process of that type. Hereinafter, a stability evaluation value indicating the stability of a process will be referred to as a process stability evaluation value. For example, in the example of FIG. 7, a process stability evaluation value is calculated for each of the process types with identifiers 1 to 13.
[0143] The monitoring function unit 72 then identifies a stability evaluation value indicating the stability of the application based on the process stability evaluation value of the process related to each type of process included in the application, which is acquired for each type of process. Hereinafter, the stability evaluation value indicating the stability of an application will be referred to as the application stability evaluation value. For example, the monitoring function unit 72 calculates the application stability evaluation value of each application based on the process stability evaluation values calculated for the processes included in the application.
[0144] Here, the application stability evaluation value may be determined based on at least one of the following: the state of a process included in the application; the lifetime of a process included in the application; the length of time that a process included in the application has performed input / output; or the number of packet drops of a process included in the application. Here, for example, the lifetime of a process can be determined based on a value indicating the start time of the process.
[0145] Alternatively, the monitoring function unit 72 may calculate an application stability evaluation value indicating the stability of an application in accordance with a rule associated with the type of application. For example, a formula may be defined in advance for each type of application. Then, the application stability evaluation value of the application may be calculated by applying the process stability evaluation value of the process related to each type, which is obtained for each type of process included in the application, to the formula.
[0146] For example, an application stability evaluation value for an application with an identifier AP1 is calculated based on the process stability evaluation values of processes of the type having identifiers 1 to 3. Furthermore, an application stability evaluation value for an application with an identifier AP2 is calculated based on the process stability evaluation values of processes of the type having identifiers 4 to 7. Furthermore, an application stability evaluation value for an application with an identifier AP3 is calculated based on the process stability evaluation values of processes of the type having identifiers 8 to 10. Furthermore, an application stability evaluation value for an application with an identifier AP4 is calculated based on the process stability evaluation values of processes of the type having identifiers 11 to 13.
[0147] The monitoring function unit 72 then generates stability evaluation value data indicating the application stability evaluation value calculated for each of the multiple applications, and outputs the generated stability evaluation value data to the data bus unit 68. In this embodiment, for example, the monitoring function unit 72 generates stability evaluation value data at predetermined time intervals based on the latest situation. The monitoring function unit 72 then outputs the stability evaluation value data to the data bus unit 68 every time the stability evaluation value data is generated.
[0148] Then, in response to the stability evaluation value data being output to the data bus unit 68, the policy manager unit 90 acquires the stability evaluation value data. The policy manager unit 90 then identifies the application stability evaluation value indicated by the acquired stability evaluation value data. In this way, the policy manager unit 90 identifies, for each of a plurality of applications, a stability evaluation value indicating the stability of the application. Furthermore, as described above, the processes included in these applications are distributed and run on a plurality of hardware resources.
[0149] The policy manager unit 90 then determines whether each of the multiple applications is unstable based on a stability evaluation value that indicates the stability of the application. For example, the more unstable the application, the smaller the application stability evaluation value. In this case, the policy manager unit 90 determines that the application is unstable when, for example, the application stability evaluation value is smaller than a threshold value associated with the type of application.
[0150] For example, suppose that the policy manager unit 90 determines that a first application (e.g., an application with an identifier AP1) is unstable. In this case, the policy manager unit 90 identifies a stability evaluation value indicating the stability of a second application having at least one process running in a hardware resource on which at least one process included in the first application is running.
[0151] In the example of FIG. 7, the servers on which processes included in the first application are running are three servers with identifiers S1, S2, and S3.
[0152] In addition to a process included in an application with an identifier AP1, a process included in an application with an identifier AP3 and a process included in an application with an identifier AP4 are running on a server with an identifier S1.
[0153] Furthermore, in the server with the identifier S2, in addition to the process included in the application with the identifier AP1, a process included in the application with the identifier AP3 is running.
[0154] Furthermore, on the server with identifier S3, in addition to the process included in the application with identifier AP1, a process included in the application with identifier AP2 and a process included in the application with identifier AP3 are running.
[0155] Therefore, in this case, the three applications with identifiers AP2, AP3, and AP4 correspond to the above-mentioned second application. In this way, there may be a plurality of second applications.
[0156] Therefore, in this case, the policy manager unit 90 identifies the application stability evaluation value for each of the three applications with identifiers AP2, AP3, and AP4.
[0157] When the policy manager unit 90 determines that the first application is unstable based on a stability evaluation value indicating the stability of the first application, it estimates, based on a stability evaluation value indicating the stability of the second application, whether the instability of the first application is caused by the first application itself or by the hardware resources on which the processes included in the first application are running.
[0158] Here, when the policy manager unit 90 determines that the second application is not unstable based on a stability evaluation value indicating the stability of the second application, the policy manager unit 90 may infer that the first application is the cause of the instability of the first application.
[0159] For example, if an application with an identifier AP2 is determined to be not unstable, it may be inferred that the cause of instability of the application with an identifier AP1 is that application. Alternatively, if an application with an identifier AP3 is determined to be not unstable, it may be inferred that the cause of instability of the application with an identifier AP1 is that application. Alternatively, if an application with an identifier AP4 is determined to be not unstable, it may be inferred that the cause of instability of the application with an identifier AP1 is that application.
[0160] In addition, when the policy manager unit 90 determines that the second application is unstable based on a stability evaluation value indicating the stability of the second application, it may infer that the instability of the first application is caused by the hardware resources on which the processes included in the first application are running.
[0161] For example, when an application with an identifier AP2 is determined to be unstable, it may be inferred that the instability of the application with an identifier AP1 is caused by the server with an identifier S3. Alternatively, when an application with an identifier AP3 is determined to be unstable, it may be inferred that the instability of the application with an identifier AP1 is caused by the server with an identifier S1, S2, or S3. Alternatively, when an application with an identifier AP4 is determined to be not unstable, it may be inferred that the instability of the application with an identifier AP1 is caused by the server with an identifier S1.
[0162] In addition, the policy manager unit 90 may estimate whether the instability of a first application is caused by the first application or the hardware resource, based on at least one of the number of applications that are determined to be unstable or the number of applications that are determined to be not unstable, among multiple applications running on any of the hardware resources on which a process included in the first application is running.
[0163] In this case, the policy manager unit 90 may infer that the reason the first application is unstable is that the hardware resource has a predetermined number of applications determined to be unstable, which is two or more.
[0164] For example, assume that the above-mentioned predetermined number is 3. In this case, if it is determined that three applications with identifiers AP1, AP3, and AP4 are unstable, it may be inferred that the cause of the instability of the first application lies in the server with identifier S1. Also, if it is determined that three applications with identifiers AP1, AP2, and AP3 are unstable, it may be inferred that the cause of the instability of the first application lies in the server with identifier S3. And, if none of the above cases occur, it may be inferred that the cause of the instability of the first application lies in that application.
[0165] Alternatively, the policy manager unit 90 may estimate that the instability of the first application is caused by a hardware resource where the ratio of the number of applications determined to be unstable to the number of applications running on the hardware resource is equal to or greater than a predetermined value.
[0166] For example, assume that the predetermined value is 60%. In this case, if 60% or more of the applications running on a server with an identifier S1 are determined to be unstable, it may be inferred that the instability of the first application is caused by the server with an identifier S1. Alternatively, if 60% or more of the applications running on a server with an identifier S2 are determined to be unstable, it may be inferred that the instability of the first application is caused by the server with an identifier S2. Alternatively, if 60% or more of the applications running on a server with an identifier S3 are determined to be unstable, it may be inferred that the instability of the first application is caused by the server with an identifier S3. Then, if neither of the above cases occurs, it may be inferred that the instability of the first application is caused by that application.
[0167] In addition, when the policy manager unit 90 determines that all applications running on any of the hardware resources on which processes included in the first application are running are unstable, the policy manager unit 90 may infer that the instability of the first application is caused by that hardware resource.
[0168] For example, if it is determined that all applications (three applications with identifiers AP1, AP3, and AP4) whose processes are running on a server with an identifier S1 are unstable, it may be inferred that the instability of the first application is caused by the server with an identifier S1. Alternatively, if it is determined that all applications (two applications with identifiers AP1 and AP3) whose processes are running on a server with an identifier S2 are unstable, it may be inferred that the instability of the first application is caused by the server with an identifier S2. Alternatively, if it is determined that all applications (three applications with identifiers AP1, AP2, and AP3) whose processes are running on a server with an identifier S3 are unstable, it may be inferred that the instability of the first application is caused by the server with an identifier S3. If neither of the above cases occurs, it may be inferred that the instability of the first application is caused by that application.
[0169] In this case, if it is determined that the three applications with identifiers AP1, AP3, and AP4 are unstable, it may be inferred that the instability of the first application is caused by the server with identifier S1 or S2. Also, if it is determined that the three applications with identifiers AP1, AP2, and AP3 are unstable, it may be inferred that the instability of the first application is caused by the server with identifier S2 or S3.
[0170] Furthermore, if it is determined that three applications with identifiers AP1, AP3, and AP4 are unstable, it may be inferred that the instability of the first application is caused by the server with identifier S1 and the server with identifier S2. Furthermore, if it is determined that three applications with identifiers AP1, AP2, and AP3 are unstable, it may be inferred that the instability of the first application is caused by the server with identifier S2 and the server with identifier S3.
[0171] In this way, it is estimated whether the cause of instability in the first application lies in the first application itself or in the hardware resources on which the processes included in the first application are running.
[0172] The policy manager unit 90 then executes an action according to the estimated cause.
[0173] Here, the policy manager unit 90 may execute replacement of the first application. For example, if it is estimated that the instability of an application with an identifier AP1 is caused by that application, the application may be replaced with a server in another cluster or with another server in the same cluster. Here, for example, the tenant settings of the application may be changed.
[0174] The policy manager unit 90 may also separate hardware resources from a cluster created by virtualization technology. For example, if it is estimated that the instability of an application with an identifier AP1 is caused by a server with an identifier S1, the server may be separated from the cluster to which it belongs.
[0175] For example, as shown in FIG. 8, a new server with an identifier S5 may be added to a cluster to which servers with identifiers S1 to S4 belong. Then, the tenant settings for AP1, AP3, and AP4 may be changed. For example, a server with an identifier S1 may be excluded from the tenant, and a server with an identifier S5 may be added to the tenant. Then, the server with an identifier S1 may be separated from the cluster to which it belongs. In this way, as shown in FIG. 8, the container management unit 78 runs a process on the server with an identifier S5 as needed.
[0176] In this embodiment, the monitoring function unit 72 may calculate the stability evaluation value of a cluster based on the stability evaluation value of an application running in that cluster.
[0177] The policy manager unit 90 may then determine whether the cluster is unstable based on the stability evaluation value of the cluster. If the cluster is determined to be unstable, the policy manager unit 90 may replace all applications running in the cluster with other clusters.
[0178] Here, an example of the flow of processing related to estimation of the cause of application instability, which is performed in the platform system 30 according to this embodiment, will be described with reference to the flow diagram shown in FIG.
[0179] In this processing example, for example, the policy manager unit 90 monitors whether stability evaluation value data indicating the stability of the application is output to the data bus unit 68 (S101).
[0180] When it is detected that the stability evaluation value data has been output to the data bus unit 68, the policy manager unit 90 acquires the stability evaluation value data (S102).
[0181] Then, the policy manager unit 90 determines whether or not the application is unstable based on the stability evaluation value data indicating the stability of the application, which is acquired in the process shown in S102 (S103).
[0182] If it is not determined to be unstable (S103: N), the process returns to S101.
[0183] If it is determined that the application is unstable (S103: Y), the policy manager unit 90 identifies the multiple hardware resources on which the processes included in the application are running (S104). In this processing example, the multiple hardware resources on which the processes included in the application are running can be identified by referring to the inventory data.
[0184] Then, the policy manager unit 90 identifies an application whose process is running on at least one of the plurality of hardware resources identified in the process shown in S104 (S105). In this process example, by referring to the inventory data, it is possible to identify, for each of the plurality of hardware resources, an application whose process is running on that hardware resource.
[0185] Then, the policy manager unit 90 identifies the latest application stability evaluation value of at least one application identified in the process shown in S105 (S106).
[0186] Then, based on the application stability evaluation value identified in the process shown in S106, the cause of the application determined to be unstable in the process shown in S103 is estimated (S107).
[0187] The policy manager unit 90 then executes an action according to the cause estimated in the process shown in S107 (S108). In the process shown in S108, for example, the policy manager unit 90, the life cycle management unit 94, the container management unit 78, and the configuration management unit 76 may cooperate with one another to execute the action. Then, the process returns to the process shown in S101.
[0188] In this embodiment, when an application is determined to be unstable based on the application stability evaluation value aggregated for multiple processes, the cause of the instability of the application is estimated based on the application stability evaluation values of other applications.
[0189] Therefore, even when the stability evaluation value of an application is identified based on the application stability evaluation values collected for a plurality of processes, it is possible to accurately estimate the cause of the instability of the application.
[0190] The present invention is not limited to the above-described embodiment.
[0191] For example, the functional units according to this embodiment are not limited to those shown in FIG.
[0192] Furthermore, the functional unit according to this embodiment does not need to be a 5G NF. For example, the functional unit according to this embodiment may be a 4G network node such as an eNodeB, a vDU, a vCU, a Packet Data Network Gateway (P-GW), a Serving Gateway (S-GW), a Mobility Management Entity (MME), or a Home Subscriber Server (HSS).
[0193] Furthermore, the scope of application of the present invention is not limited to applications included in the communication system 1. The present invention is also applicable to general applications other than those included in the communication system 1.
[0194] Furthermore, the functional units according to the present embodiment may be realized using hypervisor-type or host-type virtualization technology instead of container-type virtualization technology. Furthermore, the functional units according to the present embodiment do not need to be implemented by software, but may be implemented by hardware such as electronic circuits. Furthermore, the functional units according to the present embodiment may be implemented by a combination of electronic circuits and software.
[0195] The technology described in this disclosure can also be expressed as follows. [1] a first stability evaluation value specifying means for specifying a stability evaluation value indicating the stability of a first application whose processes are distributed and running on a plurality of hardware resources; a second stability evaluation value specifying means for specifying a stability evaluation value indicating the stability of a second application having at least one process running on the hardware resource on which at least one process included in the first application is running; an instability determination means for determining whether the application is unstable based on a stability evaluation value indicating the stability of the application; a cause estimation means for estimating, when it is determined that the first application is unstable based on a stability evaluation value indicating the stability of the first application, whether the instability of the first application is caused by the first application or by a hardware resource on which a process included in the first application is running, based on a stability evaluation value indicating the stability of the second application; A cause estimation system comprising: [2] the cause estimation means estimates that the instability of the first application is caused by the first application when it is determined that the second application is not unstable based on a stability evaluation value indicating the stability of the second application; The cause estimation system according to [1]. [3] when it is determined that the second application is unstable based on a stability evaluation value indicating the stability of the second application, the cause estimation means estimates that the instability of the first application is caused by a hardware resource on which a process included in the first application is running. The cause estimation system according to [1] or [2], [4] the cause estimation means estimates whether the cause is in the first application or in the hardware resource based on at least one of the number of applications determined to be unstable and the number of applications determined to be not unstable among the plurality of applications running on any of the hardware resources; The cause estimation system according to [1]. [5] the cause estimation means estimates that the cause is in the hardware resource where the number of applications determined to be unstable is equal to or greater than a predetermined number, which is equal to or greater than two; The cause estimation system according to [4]. [6] the cause estimation means estimates that the cause is in the hardware resource for which a ratio of the number of applications determined to be unstable to the number of applications running on the hardware resource is equal to or greater than a predetermined value. The cause estimation system according to [4]. [7] When all applications running on any one of the hardware resources are determined to be unstable, the cause estimation means estimates that the cause is in that hardware resource. The cause estimation system according to [1]. [8] the stability evaluation value indicating the stability of the application is determined based on a value indicating the stability of a process related to the type, which is obtained for each type of process included in the application; The cause estimation system according to any one of [1] to [7], [9] The process is an execution unit of the application in a container-type virtualized application execution environment. The cause estimation system according to any one of [1] to [8],
[10] The stability evaluation value indicating the stability of the application is determined based on at least one of the state of a process included in the application, the lifetime of a process included in the application, the length of time that the process included in the application has performed input / output, or the number of packet drops of the process included in the application. The cause estimation system according to any one of [1] to [9],
[11] The stability evaluation value indicating the stability of the application is calculated according to a rule associated with the type of the application. The cause estimation system according to any one of [1] to
[10] ,
[12] and an action execution means for executing an action according to the estimated cause. The cause estimation system according to any one of [1] to
[11] , characterized in that
[13] the action execution means executes replacement of the first application. The cause estimation system according to
[12] .
[14] the action execution means executes separation of the hardware resource from a cluster generated by a virtualization technology. The cause estimation system according to
[12] or
[13] ,
[15] The application is an application included in a communication system. The cause estimation system according to any one of [1] to
[14] ,
[16] The application is a network function. The cause estimation system according to
[15] .
[17] identifying a stability evaluation value indicative of stability of a first application whose processes are distributed across multiple hardware resources; identifying a stability evaluation value indicating stability of a second application having at least one process running on the hardware resource on which at least one process included in the first application is running; determining whether the application is unstable based on a stability evaluation value indicating the stability of the application; when it is determined that the first application is unstable based on a stability evaluation value indicating the stability of the first application, estimating, based on a stability evaluation value indicating the stability of the second application, whether the instability of the first application is caused by the first application itself or by a hardware resource on which a process included in the first application is running; A cause estimation method comprising:
Claims
1. a first stability evaluation value specifying means for specifying a first stability evaluation value indicating the stability of a first application whose processes are distributed and running on a plurality of hardware resources; a second stability evaluation value specifying means for specifying a second stability evaluation value indicating stability of a second application having at least one process running on the hardware resource on which at least one process included in the first application is running; an instability determination means for determining whether the first application is unstable based on the first stability evaluation value and determining whether the second application is unstable based on the second stability evaluation value; a cause estimation means for estimating, when it is determined that the first application is unstable based on the first stability evaluation value, whether the cause of instability of the first application is the first application itself or a hardware resource on which a process included in the first application is running, based on the second stability evaluation value; the cause estimation means, when it is determined that the second application is not unstable based on the second stability evaluation value, estimates that the first application is the cause of instability of the first application; Cause estimation system.
2. When it is determined that the second application is unstable based on the second stability evaluation value, the cause estimation means estimates that the instability of the first application is caused by a hardware resource on which a process included in the first application is running. The cause estimation system according to claim 1 .
3. the cause estimation means estimates whether the cause is in the first application or the hardware resource based on at least one of the number of applications determined to be unstable and the number of applications determined to be not unstable among a plurality of applications running on any of the hardware resources; The cause estimation system according to claim 1 .
4. the cause estimation means estimates that the cause is the hardware resource in which the number of applications determined to be unstable is equal to or greater than a predetermined number, which is equal to or greater than two; The cause estimation system according to claim 3 .
5. the cause estimation means estimates that the cause is in the hardware resource for which a ratio of the number of applications determined to be unstable to the number of applications running on the hardware resource is equal to or greater than a predetermined value. The cause estimation system according to claim 3 .
6. When all applications running on any one of the hardware resources are determined to be unstable, the cause estimation means estimates that the cause is in that hardware resource. The cause estimation system according to claim 1 .
7. the first stability evaluation value is determined based on a value indicating the stability of a process related to each type of process included in the first application, the value being acquired for each type of process; the second stability evaluation value is identified based on a value indicating the stability of a process related to the type, the value being acquired for each type of process included in the second application; The cause estimation system according to claim 1 .
8. The process is an execution unit of an application in a container-type virtualized application execution environment. The cause estimation system according to claim 1 .
9. the first stability evaluation value is determined based on at least one of a state of a process included in the first application, a lifetime of the process included in the first application, a length of time that the process included in the first application has performed input / output, or a number of packet drops of the process included in the first application; the second stability evaluation value is determined based on at least one of a state of a process included in the second application, a lifetime of the process included in the second application, a length of time that the process included in the second application has performed input / output, or a number of packet drops of the process included in the second application; The cause estimation system according to claim 1 .
10. the first stability evaluation value is calculated according to a rule associated with the type of the first application; the second stability evaluation value is calculated according to a rule associated with the type of the second application; The cause estimation system according to claim 1 .
11. and an action execution means for executing an action according to the estimated cause. The cause estimation system according to claim 1 .
12. the action execution means executes replacement of the first application. The cause estimation system according to claim 11.
13. the action execution means executes separation of the hardware resource from a cluster generated by a virtualization technology. The cause estimation system according to claim 11.
14. the first application and the second application are applications included in a communication system; The cause estimation system according to claim 1 .
15. the first application and the second application are network functions; The cause estimation system according to claim 14.
16. identifying a first stability metric indicating stability of a first application whose processes are distributed across multiple hardware resources; identifying a second stability evaluation value indicating stability of a second application having at least one process running on the hardware resource on which at least one process included in the first application is running; determining whether the first application is unstable based on the first stability evaluation value, and determining whether the second application is unstable based on the second stability evaluation value; when it is determined that the first application is unstable based on the first stability evaluation value, estimating, based on the second stability evaluation value, whether the instability of the first application is caused by the first application itself or by a hardware resource on which a process included in the first application is running; In the estimation, when it is determined that the second application is not unstable based on the second stability evaluation value, it is estimated that the instability of the first application is caused by the first application. Cause estimation method.
Citation Information
Patent Citations
Method and apparatus for eliminating a single point of failure in cloud-based applications
JP2015522876A
Diagnostic program, diagnostic method, and diagnostic apparatus
JP2019012477A
Control program, control method, and control apparatus
JP2021144401A
System and method for monitoring an application or service group within a cluster as a resource of another cluster
US8464092B1
Resource management method and resource management system
WO2015145664A1