Determining whether a process included in a communication system is unstable.
The determination system addresses the inefficiency of monitoring distributed processes by using a platform system with advanced components to identify unstable processes across virtual machines, enhancing monitoring efficiency and accuracy.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-12-26
- Publication Date
- 2026-03-31
AI Technical Summary
Monitoring the stability of distributed processes across multiple virtual machines in a communication system is resource-intensive and inefficient in existing systems.
A determination system that includes application monitoring and process instability determination means to identify unstable processes across multiple virtual machines, utilizing a platform system with components like a monitoring function unit, AI unit, and data bus unit to aggregate and analyze performance and stability data.
Efficiently identifies unstable processes, reducing resource overhead and improving monitoring accuracy in communication systems with distributed processes.
Smart Images

Figure 0007838123000001 
Figure 0007838123000002 
Figure 0007838123000003
Abstract
Description
Technical Field
[0001] The present invention relates to determining whether a process included in a communication system is unstable or not.
Background Art
[0002] Patent Document 1 describes NFV (Network Functions Virtualization) that realizes functions such as network devices software-wise by a virtual machine (VM) implemented on a virtualization layer such as a hypervisor.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004]
[0005] Among applications such as network functions operating in a communication system, there are some that include a plurality of processes, and these processes are distributed and operating on a plurality of virtual machines.
[0006] [[ID=Finally, when extracting an unstable process from these processes, if each process is monitored for whether it has become unstable, the processing load for monitoring may become extremely large.
Means for Solving the Problems
[0007] To solve the above problems, the determination system according to this disclosure includes application monitoring means for monitoring whether an application included in a communication system, in which the processes included are running distributed across multiple virtual machines, has become unstable, and process instability determination means for determining whether the processes running on each of the multiple virtual machines in which at least one process included in the application is running are unstable, in response to the detection that the application has become unstable.
[0008] Furthermore, the determination method relating to this disclosure includes monitoring whether an application included in a communication system, whose processes are distributed and run across multiple virtual machines, has become unstable, and, in response to the detection of instability in the said application, determining whether the processes running on each of the multiple virtual machines on which at least one process included in the said application is running are unstable. [Brief explanation of the drawing]
[0009] [Figure 1] This figure shows an example of a communication system according to one embodiment of the present invention. [Figure 2] This figure shows an example of a communication system according to one embodiment of the present invention. [Figure 3] This diagram schematically illustrates an example of a network service related to one embodiment of the present invention. [Figure 4] This figure shows an example of the relationships between elements constructed in a communication system according to one embodiment of the present invention. [Figure 5] This is a functional block diagram showing an example of a function implemented in a platform system according to one embodiment of the present invention. [Figure 6] This figure shows an example of the data structure of physical inventory data. [Figure 7] This diagram schematically illustrates an example of a situation where processes included in multiple applications are distributed and run across multiple virtual machines. [Figure 8] This diagram schematically illustrates an example of a situation where processes included in multiple applications are distributed and run across multiple virtual machines. [Figure 9] This diagram schematically illustrates an example of a situation where processes included in multiple applications are distributed and run across multiple virtual machines. [Figure 10A] This flowchart shows an example of the processing flow performed in a platform system related to a certain model. [Figure 10B] This flowchart shows an example of the processing flow performed in a platform system related to a certain model. [Modes for carrying out the invention]
[0010] One embodiment of the present invention will be described in detail below with reference to the drawings.
[0011] Figures 1 and 2 show an example of a communication system 1 according to one embodiment of the present invention. Figure 1 focuses on the location of the data center group included in the communication system 1. Figure 2 focuses on the various computer systems implemented in the data center group included in the communication system 1.
[0012] As shown in Figure 1, the data center group included in communication system 1 is classified into a central data center 10, a regional data center 12, and an edge data center 14.
[0013] The central data center 10 is distributed among several locations within the area covered by the communication system 1 (for example, within Japan).
[0014] The regional data centers 12 are, for example, distributed among dozens of others within the area covered by the communication system 1. For example, if the area covered by the communication system 1 is the entire country of Japan, then one or two regional data centers 12 may be located in each prefecture.
[0015] The edge data centers 14 are, for example, distributed and arranged in thousands within the area covered by the communication system 1. Each of the edge data centers 14 can communicate with communication facilities 18 equipped with antennas 16. As shown in FIG. 1 here, one edge data center 14 may be able to communicate with several communication facilities 18. The communication facilities 18 may include computers such as server computers. The communication facilities 18 according to the present embodiment perform wireless communication with a UE (User Equipment) 20 via the antenna 16. For example, an RU (Radio Unit) described later is provided in the communication facilities 18 equipped with the antenna 16.
[0016] In the central data center 10, regional data center 12, and edge data center 14 according to the present embodiment, a plurality of servers are respectively arranged.
[0017] In the present embodiment, for example, the central data center 10, regional data center 12, and edge data center 14 can communicate with each other. Also, the central data centers 10 can communicate with each other, the regional data centers 12 can communicate with each other, and the edge data centers 14 can communicate with each other.
[0018] As shown in FIG. 2, the communication system 1 according to the present embodiment includes a platform system 30, a plurality of radio access networks (RANs) 32, a plurality of core network systems 34, and a plurality of UEs 20. The core network system 34, RAN 32, and UE 20 cooperate with each other to realize a mobile communication network.
[0019] RAN32 is a computer system equipped with an antenna 16, equivalent to an eNB (eNodeB) in a fourth-generation mobile communication system (hereinafter referred to as 4G) or a gNB (NR base station) in a fifth-generation mobile communication system (hereinafter referred to as 5G). In this embodiment, RAN32 is mainly implemented by a group of servers and communication equipment 18 located in an edge data center 14. However, some parts of RAN32 (for example, DU (Distributed Unit), CU (Central Unit), vDU (virtual Distributed Unit), vCU (virtual Central Unit)) may be implemented in a central data center 10 or a regional data center 12 instead of the edge data center 14.
[0020] The core network system 34 is a system equivalent to the EPC (Evolved Packet Core) in 4G or the 5G core (5GC) in 5G. The core network system 34 according to this embodiment is mainly implemented by a group of servers located in the central data center 10 and the regional data center 12.
[0021] The platform system 30 according to this embodiment is configured, for example, on a cloud infrastructure and includes a processor 30a, a storage unit 30b, and a communication unit 30c, as shown in Figure 2. The processor 30a is a program control device such as a microprocessor that operates according to a program installed on the platform system 30. The storage unit 30b is, for example, a memory element such as ROM or RAM, or a solid-state drive (SSD) or hard disk drive (HDD). The storage unit 30b stores programs executed by the processor 30a. The communication unit 30c is, for example, a communication interface such as a NIC (Network Interface Controller) or a wireless LAN (Local Area Network) module. Software-Defined Networking (SDN) may be implemented in the communication unit 30c. The communication unit 30c exchanges data with the RAN 32 and the core network system 34.
[0022] In this embodiment, the platform system 30 is implemented by a group of servers located in the central data center 10. Alternatively, the platform system 30 may be implemented by a group of servers located in the regional data center 12.
[0023] In this embodiment, for example, in response to a purchase request for network services (NS) from a purchaser, the requested network services are built in RAN32 or the core network system 34. The built network services are then provided to the purchaser.
[0024] For example, a purchaser, who is an MVNO (Mobile Virtual Network Operator), is provided with network services such as voice communication services and data communication services. The voice communication services and data communication services provided by this embodiment are ultimately provided to the customer (end user) of the purchaser (MVNO in the above example) who uses the UE20 shown in Figures 1 and 2. This end user can perform voice and data communication with other users via the RAN32 and the core network system 34. Furthermore, the end user's UE20 is able to access data networks such as the Internet via the RAN32 and the core network system 34.
[0025] Furthermore, in this embodiment, IoT (Internet of Things) services may be provided to end users who utilize robotic arms, connected cars, etc. In this case, for example, the end users who utilize robotic arms, connected cars, etc. may become purchasers of the network services according to this embodiment.
[0026] In this embodiment, the servers located in the central data center 10, the regional data center 12, and the edge data center 14 have a hypervisor (bare-metal hypervisor) and host-based virtualization software running on a host operating system (host OS) installed. One or more virtual machines (VMs) run on each server. One or more processes can be deployed and run on each virtual machine. In this embodiment, a cluster of virtual machines spanning multiple servers may be constructed.
[0027] In this embodiment, the network service provided to the purchaser consists of one or more functional units (e.g., network functions (NFs)). In this embodiment, the functional unit is implemented using an NF realized by virtualization technology. An NF realized by virtualization technology is referred to as a VNF (Virtualized Network Function). In the following description, the functional unit is assumed to be implemented using a VNF realized by hypervisor-type or host-type virtualization technology. In this embodiment, the network service is described as being implemented by one or more NFs. Furthermore, the functional unit in this embodiment may correspond to a network node.
[0028] Figure 3 is a schematic diagram illustrating an example of a network service in operation. The network service shown in Figure 3 includes several RU40s, several DU42s, several CU44s (CU-CP (Central Unit - Control Plane) 44a and CU-UP (Central Unit - User Plane) 44b), several AMFs (Access and Mobility Management Functions) 46, several SMFs (Session Management Functions) 48, and several UPFs (User Plane Functions) 50 as software elements.
[0029] In the example in Figure 3, RU40, DU42, CU-CP44a, AMF46, and SMF48 correspond to elements of the control plane (C-Plane), while RU40, DU42, CU-UP44b, and UPF50 correspond to elements of the user plane (U-Plane).
[0030] Furthermore, the network service may include other types of network infrastructure (NF) as software elements. Also, the network service is implemented on multiple computer resources (hardware elements) such as servers.
[0031] In this embodiment, for example, a communication service in a certain area is provided by the network service shown in Figure 3.
[0032] In this embodiment, the multiple RU40s, multiple DU42s, multiple CU-UP44bs, and multiple UPF50s shown in Figure 3 belong to a single end-to-end network slice.
[0033] Figure 4 is a schematic diagram illustrating an example of the relationships between elements constructed in the communication system 1 in this embodiment. The symbols M and N shown in Figure 4 represent any integer of 1 or more, indicating the relationship between the number of elements connected by a link. When both ends of a link are a combination of M and N, the elements connected by that link have a many-to-many relationship. When both ends of a link are a combination of 1 and N or 1 and M, the elements connected by that link have a one-to-many relationship.
[0034] As shown in Figure 4, network services (NS), network functions (NF), and processes are arranged in a hierarchical structure.
[0035] NS corresponds to, for example, a network service composed of multiple NFs. Here, NS may correspond to elements at a granularity such as 5GC, EPC, 5G RAN (gNB), 4G RAN (eNB), etc.
[0036] In 5G, NFs correspond to elements of a granularity such as RU, DU, CU-CP, CU-UP, AMF, SMF, and UPF. In 4G, NFs correspond to elements of a granularity such as MME (Mobility Management Entity), HSS (Home Subscriber Server), S-GW (Serving Gateway), vDU, and vCU. In this embodiment, for example, one NS contains one or more NFs. That is, one or more NFs are under the control of one NS.
[0037] Furthermore, an NF (Field Process) contains one or more processes. In other words, one or more processes are under the control of a single NF.
[0038] A process may provide some of the functions of DU, CU-CP, CU-UP, etc. Similarly, a process may provide some of the functions of UPF, AMF, SMF, etc. For example, a UPF may include multiple types of processes, such as a management process and a user plane communication process. Furthermore, a single NF (e.g., a single UPF) may contain multiple processes of a specific type.
[0039] Furthermore, as shown in Figure 4, network slices (NSIs) and network slice subnet instances (NSSIs) are arranged in a hierarchical structure.
[0040] An NSI can be described as an end-to-end virtual circuit spanning multiple domains (for example, from RAN32 to core network system34). An NSI may be a slice for high-speed, high-capacity communication (e.g., for eMBB: enhanced Mobile Broadband), a slice for highly reliable and low-latency communication (e.g., for URLLC: Ultra-Reliable and Low Latency Communications), or a slice for connecting a large number of terminals (e.g., for mMTC: massive Machine Type Communication). An NSSI can also be described as a virtual circuit in a single domain obtained by dividing an NSI. An NSSI may be a slice in the RAN domain, a slice in a transport domain such as the MBH (Mobile Back Haul) domain, or a slice in the core network domain.
[0041] In this embodiment, for example, one NSI contains one or more NSSIs. That is, one or more NSSIs are under the control of one NSI. In this embodiment, multiple NSIs may share the same NSSI.
[0042] Furthermore, as shown in Figure 4, NSSI and NS generally have a many-to-many relationship.
[0043] Furthermore, in this embodiment, for example, one NF can belong to one or more network slices. Specifically, for example, one NF can be configured with NSSAI (Network Slice Selection Assistance Information) that includes one or more S-NSSAI (Sub Network Slice Selection Assist Information). Here, S-NSSAI is information associated with a network slice. Note that an NF does not necessarily have to belong to a network slice.
[0044] Figure 5 is a functional block diagram showing an example of the functions implemented in the platform system 30 according to this embodiment. Note that not all of the functions shown in Figure 5 are required to be implemented in the platform system 30 according to this embodiment, and other functions may also be implemented.
[0045] As shown in Figure 5, the platform system 30 according to this embodiment functionally includes, for example, an Operation Support System (OSS) unit 60, an Orchestration (E2EO: End-to-End-Orchestration) unit 62, a Service Catalog Storage Unit 64, a Big Data Platform Unit 66, a Data Bus Unit 68, an Artificial Intelligence (AI) unit 70, a Monitoring Function Unit 72, an SDN Controller 74, a Configuration Management Unit 76, a Process Management Unit 78, and a Repository Unit 80. The OSS unit 60 includes an Inventory Database 82, a Ticket Management Unit 84, a Fault Management Unit 86, and a Performance Management Unit 88. The E2EO unit 62 includes a Policy Manager Unit 90, a Slice Manager Unit 92, and a Lifecycle Management Unit 94. These elements are mainly implemented as a processor 30a, a storage unit 30b, and a communication unit 30c.
[0046] The functions shown in Figure 5 may be implemented by installing them on a platform system 30, which is one or more computers, and having a processor 30a execute a program containing commands corresponding to those functions. This program may be supplied to the platform system 30 via a computer-readable information storage medium such as an optical disk, magnetic disk, magnetic tape, magneto-optical disk, or flash memory, or via the internet. The functions shown in Figure 5 may also be implemented using circuit blocks, memory, or other LSIs. Furthermore, it will be understood by those skilled in the art that the functions shown in Figure 5 can be realized in various forms, such as hardware only, software only, or a combination thereof.
[0047] The process management unit 78 performs process lifecycle management. For example, processes related to process construction, such as process deployment and configuration, are included in this lifecycle management.
[0048] Here, the platform system 30 according to this embodiment may include a plurality of process management units 78. A process management tool may be installed on each of the plurality of process management units 78. Each of the plurality of process management units 78 may perform process construction, such as process deployment, on a group of servers (e.g., a cluster) associated with the process management unit 78.
[0049] The process control unit 78 does not need to be included in the platform system 30. The process control unit 78 may, for example, be located on a server managed by the process control unit 78 (i.e., RAN 32 or the core network system 34), or on another server co-located with the server managed by the process control unit 78.
[0050] In this embodiment, the repository unit 80 stores, for example, images of processes included in a group of functional units (e.g., an NF group) that implement network services.
[0051] The inventory database 82 is a database that stores inventory information. This inventory information includes, for example, information about servers located in RAN32 and the core network system 34 and managed by the platform system 30.
[0052] In this embodiment, the inventory database 82 stores inventory data. The inventory data shows the current configuration of the elements included in the communication system 1 and the relationships between those elements. The inventory data also shows the status of resources managed by the platform system 30 (for example, resource usage). This inventory data may be physical inventory data or logical inventory data. Physical inventory data and logical inventory data will be described later.
[0053] Figure 6 shows an example of the data structure of physical inventory data. The physical inventory data shown in Figure 6 is associated with one server. The physical inventory data shown in Figure 6 includes, for example, server ID, location data, building data, floor number data, rack data, specification data, network data, running process ID list, cluster ID, etc.
[0054] The server ID included in the physical inventory data is, for example, an identifier for the server associated with that physical inventory data.
[0055] Location data included in physical inventory data is, for example, data indicating the location (e.g., the address of the location) of the server associated with that physical inventory data.
[0056] The building data included in the physical inventory data is, for example, data indicating the building (e.g., building name) where the server associated with that physical inventory data is located.
[0057] The floor number data included in the physical inventory data is, for example, data indicating the floor on which the server associated with that physical inventory data is located.
[0058] The rack data included in the physical inventory data is, for example, an identifier for the rack where the server associated with that physical inventory data is located.
[0059] The specification data included in the physical inventory data is, for example, data that indicates the specifications of the server associated with that physical inventory data, and the specification data includes things like the number of cores, memory capacity, and hard disk capacity.
[0060] The network data included in the physical inventory data is, for example, data that shows information about the network of the server associated with the physical inventory data. The network data includes, for example, the NIC (Network Interface Card) that the server has, the number of ports that the NIC has, and the port IDs of those ports.
[0061] The list of operational process IDs included in the physical inventory data is, for example, data that shows information about one or more processes running on the server associated with the physical inventory data, and the list of operational process IDs shows, for example, a list of instance identifiers (process IDs) of the process.
[0062] The cluster ID included in the physical inventory data is, for example, the identifier of the cluster (e.g., the Cubanetes cluster) to which the server associated with that physical inventory data belongs.
[0063] The logical inventory data includes topology data showing the current state of relationships between elements, as shown in Figure 4, for multiple elements included in the communication system 1. For example, the logical inventory data includes topology data that includes the identifier of a certain NS and the identifiers of one or more NFs under that NS. Also, for example, the logical inventory data includes topology data that includes the identifier of a certain network slice and the identifiers of one or more NFs belonging to that network slice.
[0064] Furthermore, the inventory data may include data indicating the current status of geographical and topological relationships between elements included in the communication system 1. As mentioned above, the inventory data includes location data indicating the locations where the elements included in the communication system 1 are operating, that is, the current locations of the elements included in the communication system 1. From this, it can be said that the inventory data indicates the current status of geographical relationships between elements (for example, geographical proximity between elements).
[0065] Furthermore, the logical inventory data may include NSI data that indicates information about network slices. NSI data indicates attributes such as the identifier of a network slice instance and the type of network slice. Additionally, the logical inventory data may include NSSI data that indicates information about network slice subnets. NSSI data indicates attributes such as the identifier of a network slice subnet and the type of network slice subnet.
[0066] Furthermore, the logical inventory data may include NS data that indicates information about NS. NS data may, for example, indicate the identifier of an NS instance or attributes such as the type of NS. The logical inventory data may also include NF data that indicates information about NF. NF data may, for example, indicate the identifier of an NF instance or attributes such as the type of NF. The logical inventory data may also include process data that indicates information about processes included in an NF. Process data may, for example, indicate the process ID of a process instance or attributes such as the type of process.
[0067] The process ID in the process data included in the logical inventory data and the process ID in the list of running process IDs included in the physical inventory data are used to associate a process instance with the server on which that process instance is running.
[0068] Furthermore, data indicating various attributes such as hostnames and IP addresses may be included in the aforementioned data contained in the logical inventory data. For example, process data may include data indicating the IP address of the process corresponding to that process data. Also, for example, NF data may include data indicating the IP address and hostname of the NF indicated by that NF data.
[0069] Furthermore, the logical inventory data may include data indicating NSSAIs, which include one or more S-NSSAIs, that are set for each NF.
[0070] Furthermore, the inventory database 82 works in conjunction with the process management unit 78 to monitor the status of resources as needed. The inventory database 82 then updates the inventory data stored in it as needed based on the latest status of the resources.
[0071] Furthermore, in response to actions such as the construction of new elements included in communication system 1, the configuration of elements included in communication system 1, scaling of elements included in communication system 1, or replacement of elements included in communication system 1, the inventory database 82 updates the inventory data stored in the inventory database 82.
[0072] The service catalog storage unit 64 stores service catalog data. The service catalog data may include, for example, service template data that shows logic used by the lifecycle management unit 94. This service template data includes information necessary to build network services. For example, the service template data includes information that defines NS, NF and processes, and information that shows the correspondence between NS, NF and processes. Also, for example, the service template data includes a script for a workflow to build network services.
[0073] An example of service template data is an NSD (NS Descriptor). An NSD is associated with a network service and indicates the types of multiple functional units included in that network service. The NSD may also indicate the number of each type of functional unit included in the network service. Furthermore, the NSD may indicate the file name of the NFD (Non-Functional Data) related to the NFs included in the network service, as described later.
[0074] Another example of service template data is an NFD (NF Descriptor). The NFD may indicate the computer resources required by the NF (e.g., CPU, memory, hard disk, etc.). For example, the NFD may indicate the computer resources required by each of the multiple processes included in the NF (CPU, memory, hard disk, etc.).
[0075] Furthermore, the service catalog data may include information about thresholds (e.g., anomaly detection thresholds) used by the policy manager unit 90 to compare with calculated performance indicator values and stability evaluation values. Performance indicator values and stability evaluation values will be described later.
[0076] Furthermore, the service catalog data may also include, for example, slice template data. The slice template data contains information necessary to perform instantiation of network slices, and includes, for example, logic used by the slice manager unit 92.
[0077] Slice template data includes information on the "Generic Network Slice Template" defined by the GSMA (GSM Association) ("GSM" is a registered trademark). Specifically, slice template data includes network slice template data (NST), network slice subnet template data (NSST), and network service template data. Furthermore, slice template data includes information showing the hierarchical structure of these elements, as shown in Figure 4.
[0078] In this embodiment, the lifecycle management unit 94, for example, constructs a new network service in response to a purchase request for an NS from a purchaser.
[0079] The lifecycle management unit 94 may, for example, execute a workflow script associated with the network service to be purchased in response to a purchase request. By executing this workflow script, the lifecycle management unit 94 may instruct the process management unit 78 to deploy the processes included in the newly purchased network service. The process management unit 78 may then retrieve an image of the process from the repository unit 80 and deploy the process corresponding to that image to the server.
[0080] Furthermore, in this embodiment, the lifecycle management unit 94 performs scaling and replacement of elements included in the communication system 1, for example. Here, the lifecycle management unit 94 may output process deployment and deletion instructions to the process management unit 78. The process management unit 78 may then perform processes such as process deployment and process deletion in accordance with these instructions. In this embodiment, the lifecycle management unit 94 enables scaling and replacement that cannot be handled by the tools of the process management unit 78.
[0081] The lifecycle management unit 94 may also output an instruction to the SDN controller 74 to create a communication path. For example, the lifecycle management unit 94 may provide the SDN controller 74 with two IP addresses at both ends of the communication path to be created, and the SDN controller 74 will create a communication path connecting these two IP addresses. The created communication path may be managed in association with these two IP addresses.
[0082] Furthermore, the lifecycle management unit 94 may output an instruction to the SDN controller 74 to create a communication path between the two IP addresses associated with those two IP addresses.
[0083] In this embodiment, the slice manager unit 92 performs, for example, the instantiation of a network slice. In this embodiment, the slice manager unit 92 performs, for example, the instantiation of a network slice by executing the logic indicated by the slice template stored in the service catalog storage unit 64.
[0084] The slice manager unit 92 includes the functions of NSMF (Network Slice Management Function) and NSSMF (Network Slice Sub-network Management Function), as described in, for example, the 3GPP (Third Generation Partnership Project) specification "TS28 533". NSMF is a function that generates and manages network slices and provides management services for NSI. NSSMF is a function that generates and manages network slice subnets that constitute a part of a network slice and provides management services for NSSI.
[0085] Here, the slice manager unit 92 may output configuration management instructions related to the instantiation of network slices to the configuration management unit 76. The configuration management unit 76 may then perform configuration management, such as setting, in accordance with the said configuration management instructions.
[0086] The slice manager unit 92 may also present two IP addresses to the SDN controller 74 and output an instruction to create a communication path between these two IP addresses.
[0087] In this embodiment, the configuration management unit 76 performs configuration management, such as setting up groups of elements like NFs, in accordance with configuration management instructions received from, for example, the lifecycle management unit 94 or the slice manager unit 92.
[0088] In this embodiment, the SDN controller 74 creates a communication path between two IP addresses associated with a communication path creation instruction, for example, in accordance with the instruction received from the lifecycle management unit 94 or the slice manager unit 92. The SDN controller 74 may create the communication path between the two IP addresses using a known path calculation method, such as Flex Algo.
[0089] For example, the SDN controller 74 may use segment routing technology (e.g., SRv6 (Segment Routing IPv6)) to build NSIs and NSSIs for aggregation routers and servers located between communication paths. Alternatively, the SDN controller 74 may generate NSIs and NSSIs across multiple target NFs by issuing commands to configure a common VLAN (Virtual Local Area Network) for multiple target NFs, and commands to assign the bandwidth and priority indicated in the configuration information to that VLAN.
[0090] Furthermore, the SDN controller 74 may perform actions such as changing the maximum bandwidth available for communication between two IP addresses without constructing a network slice.
[0091] The platform system 30 according to this embodiment may include a plurality of SDN controllers 74. Each of the plurality of SDN controllers 74 may perform processing such as creating communication paths for a group of network devices such as AGs associated with the SDN controller 74.
[0092] In this embodiment, the monitoring function unit 72 monitors, for example, the group of elements included in the communication system 1 according to a given management policy. Here, the monitoring function unit 72 may monitor the group of elements according to a monitoring policy specified by the purchaser when purchasing the network service, for example.
[0093] In this embodiment, the monitoring function unit 72 performs monitoring at various levels, such as the slice level, NS level, NF level, process level, and hardware level such as servers.
[0094] The monitoring function unit 72 may, for example, configure modules that output metric data to hardware such as a server or software elements included in the communication system 1, so that monitoring can be performed at the various levels described above. For example, an NF may output metric data to the monitoring function unit 72 that indicates metrics that can be measured (identified) in the NF. Alternatively, a server may output metric data to the monitoring function unit 72 that indicates metrics related to hardware that can be measured (identified) in the server.
[0095] Alternatively, for example, the monitoring function unit 72 may deploy a sidecar process on the server to acquire metric data on a process-by-process basis. The monitoring function unit 72 may also repeatedly execute the process of acquiring metric data acquired on a process-by-process basis from the sidecar process at a given monitoring interval, utilizing the mechanism of the monitoring tool.
[0096] The monitoring function unit 72 may, for example, monitor performance indicator values for performance indicators described in "TS 28.552, Management and orchestration; 5G performance measurements" or "TS 28.554, Management and orchestration; 5G end to end Key Performance Indicators (KPIs)". The monitoring function unit 72 may also acquire metric data indicating the monitored performance indicator values.
[0097] In this embodiment, the monitoring function unit 72 generates performance index value data indicating the performance index values of the elements included in the communication system 1 in a predetermined aggregation unit by performing a process (enrichment) to aggregate metric data in a predetermined aggregation unit.
[0098] For example, for a single gNB, performance index data for that gNB is generated by aggregating metric data that shows the metrics of the elements under that gNB (e.g., network nodes such as DU42 and CU44). In this way, performance index data showing the communication performance in the area covered by that gNB is generated. Here, for example, performance index data showing multiple types of communication performance, such as traffic volume (throughput) and latency, may be generated for each gNB. Note that the communication performance shown by the performance index data is not limited to traffic volume or latency.
[0099] The monitoring function unit 72 then outputs the performance indicator value data generated by the enrichment described above to the data bus unit 68.
[0100] In this embodiment, the data bus unit 68 receives, for example, performance index value data output from the monitoring function unit 72. The data bus unit 68 then generates a performance index value file containing the received performance index value data (one or more). The data bus unit 68 then outputs the generated performance index value file to the big data platform unit 66.
[0101] Furthermore, in this embodiment, the monitoring function unit 72 identifies, for example, a stability evaluation value indicating the stability of the elements included in the communication system 1. The monitoring function unit 72 then generates stability evaluation value data indicating the identified stability evaluation value. The monitoring function unit 72 then outputs the generated stability evaluation value data to the data bus unit 68.
[0102] In this embodiment, the data bus unit 68 receives, for example, stability evaluation value data output from the monitoring function unit 72.
[0103] Furthermore, elements such as network slices, NS, NF, and processes included in the communication system 1, as well as hardware such as servers, notify the monitoring function unit 72 of various alerts (for example, alerts triggered by the occurrence of a failure).
[0104] Then, when the monitoring function unit 72 receives, for example, the notification of the alert mentioned above, it outputs alert message data indicating the notification to the data bus unit 68. The data bus unit 68 then generates an alert file by combining the alert message data indicating one or more notifications into a single file, and outputs the alert file to the big data platform unit 66.
[0105] In this embodiment, the big data platform unit 66 stores, for example, performance indicator value files and alert files output from the data bus unit 68.
[0106] In this embodiment, the AI unit 70 has, for example, several pre-trained machine learning models stored in it. The AI unit 70 uses the various machine learning models stored in it to perform estimation processing, such as future prediction processing of the usage status and service quality of the communication system 1. The AI unit 70 may also generate estimation result data that shows the results of the estimation processing.
[0107] The AI unit 70 may perform estimation processing based on the files stored in the big data platform unit 66 and the machine learning model described above. This estimation processing is suitable when predicting long-term trends at a low frequency.
[0108] Furthermore, the AI unit 70 is capable of acquiring performance indicator data stored in the data bus unit 68. The AI unit 70 may perform estimation processing based on the performance indicator data stored in the data bus unit 68 and the machine learning model described above. This estimation processing is suitable when short-term predictions are made frequently.
[0109] In this embodiment, the performance management unit 88 calculates performance indicator values (e.g., KPIs) based on the metrics indicated by multiple metric data, for example. The performance management unit 88 may also calculate performance indicator values that are an overall evaluation of multiple types of metrics that cannot be calculated from a single metric data (e.g., performance indicator values related to end-to-end network slices). The performance management unit 88 may also generate overall performance indicator value data that shows the performance indicator value which is an overall evaluation.
[0110] The performance management unit 88 may also obtain the performance indicator value file mentioned above from the big data platform unit 66. Furthermore, the performance management unit 88 may obtain estimated result data from the AI unit 70. Based on at least one of the performance indicator value file or the estimated result data, it may calculate performance indicator values such as KPIs. Alternatively, the performance management unit 88 may directly obtain metric data from the monitoring function unit 72. Based on this metric data, it may calculate performance indicator values such as KPIs.
[0111] In this embodiment, the fault management unit 86 detects the occurrence of a fault in the communication system 1 based on, for example, at least one of the metric data, the alert notification, the estimated result data, and the overall performance index value data described above. The fault management unit 86 may, for example, detect the occurrence of a fault that cannot be detected from a single metric data or a single alert notification based on predetermined logic. The fault management unit 86 may generate detected fault data indicating the detected fault.
[0112] Furthermore, the fault management unit 86 may directly obtain metric data and alert notifications from the monitoring function unit 72. The fault management unit 86 may also obtain performance indicator value files and alert files from the big data platform unit 66. Additionally, the fault management unit 86 may obtain alert message data from the data bus unit 68.
[0113] In this embodiment, the policy manager unit 90 performs a predetermined determination process based on at least one of the above-mentioned metric data, performance indicator value data, stability evaluation value data, alert message data, performance indicator value file, alert file, estimated result data, overall performance indicator value data, and detected failure data.
[0114] The policy manager unit 90 may then perform actions according to the result of the determination process. For example, the policy manager unit 90 may output instructions to the slice manager unit 92 to build a network slice. Alternatively, the policy manager unit 90 may output instructions to the lifecycle management unit 94 to scale or replace elements, depending on the result of the determination process.
[0115] The policy manager unit 90 according to this embodiment is capable of acquiring performance indicator value data stored in the data bus unit 68. The policy manager unit 90 may then perform a predetermined determination process based on the performance indicator value data acquired from the data bus unit 68. Alternatively, the policy manager unit 90 may also perform a predetermined determination process based on alert message data stored in the data bus unit 68.
[0116] Furthermore, the policy manager unit 90 according to this embodiment is capable of acquiring stability evaluation value data stored in the data bus unit 68. The policy manager unit 90 may then perform a predetermined determination process based on the stability evaluation value data acquired from the data bus unit 68. For example, the policy manager unit 90 may determine whether or not an application is unstable based on stability evaluation value data indicating the stability of the application.
[0117] In this embodiment, the ticket management unit 84 generates a ticket indicating the contents to be notified to the administrator of the communication system 1. The ticket management unit 84 may also generate a ticket indicating the contents of the incident data. The ticket management unit 84 may also generate a ticket indicating the values of performance indicator data, stability evaluation data, or metric data. The ticket management unit 84 may also generate a ticket indicating the judgment result by the policy manager unit 90.
[0118] The ticket management unit 84 then notifies the administrator of the communication system 1 of the generated ticket. The ticket management unit 84 may, for example, send an email with the generated ticket attached to the email address of the administrator of the communication system 1.
[0119] The platform system 30 according to this embodiment determines whether or not a process included in the communication system 1 is unstable. The determination of whether or not a process is unstable, performed by the platform system 30 according to this embodiment, will be further described below.
[0120] Figure 7 schematically illustrates an example of a situation where processes included in multiple applications are distributed and run across multiple virtual machines.
[0121] The example in Figure 7 shows a situation where four applications, each with identifiers AP1, AP2, AP3, and AP4, are running.
[0122] These applications may also be network functions (e.g., DU42, CU-CP44a, CU-UP44b, AMF46, SMF48, UPF50, etc.).
[0123] In this embodiment, the hardware resources capable of running each type of application are predetermined. In the following description, the hardware resources are assumed to be servers, but these hardware resources do not necessarily have to be servers; for example, they could be nodes.
[0124] Hereafter, a hardware resource capable of running a particular type of application will be referred to as the tenant corresponding to that application.
[0125] Figure 7 shows four servers with identifiers S1, S2, S3, and S4, respectively. These four servers belong to a single cluster with identifier CL101.
[0126] Furthermore, the four applications shown in Figure 7 are assumed to be of different types. The tenant corresponding to the application with identifier AP1 will contain servers with identifiers S1, S2, and S3. The tenant corresponding to the application with identifier AP2 will contain servers with identifiers S2 and S3. The tenant corresponding to the application with identifier AP3 will contain servers with identifiers S1, S2, S3, and S4. The tenant corresponding to the application with identifier AP4 will contain servers with identifiers S1 and S4.
[0127] Furthermore, the server with identifier S1 is assumed to be running virtual machines with identifiers VM1, VM6, and VM10. Similarly, the server with identifier S2 is assumed to be running virtual machines with identifiers VM2, VM4, and VM7. The server with identifier S3 is assumed to be running virtual machines with identifiers VM3, VM5, and VM8. And the server with identifier S4 is assumed to be running virtual machines with identifiers VM9 and VM11.
[0128] Furthermore, the rounded rectangles shown in Figure 7 represent a single process. The numbers shown on the rounded rectangles are identifiers that correspond to the type of process.
[0129] As shown in Figure 7, the application with identifier AP1 contains one of each of three types of processes with identifiers 1, 2, and 3. These processes with identifiers 1, 2, and 3 are running on virtual machines with identifiers VM1, VM2, and VM3, respectively.
[0130] The application with identifier AP2 contains one of two types of processes, one with identifier 4 and one with identifier 5. These processes with identifiers 4 and 5 are running on virtual machines with identifiers VM4 and VM5, respectively.
[0131] Furthermore, the application with identifier AP3 contains one of each of the four types of processes, each with identifiers 6, 7, 8, and 9. These processes with identifiers 6, 7, 8, and 9 are running on virtual machines with identifiers VM6, VM7, VM8, and VM9, respectively.
[0132] The application with identifier AP4 contains one of two types of processes, one each with identifiers 10 and 11. These processes with identifiers 10 and 11 are running on virtual machines with identifiers VM10 and VM11, respectively.
[0133] As shown in Figure 7, the application contains multiple processes. These processes will then run in a distributed manner across multiple virtual machines.
[0134] In the example in Figure 7, one application contains one of each type of process, but one application may contain multiple processes of the same type. Also, in the example in Figure 7, one virtual machine is running one process, but one virtual machine may be running multiple processes.
[0135] In this embodiment, for example, the monitoring function unit 72 acquires a value (metric) indicating the stability of each process. Here, for example, metrics such as a value indicating the state of the process (e.g., alive or dead), the start time of the process, the length of time the process performed input / output, the number of packets dropped by the process, and the number of packets dropped by the process may be acquired.
[0136] The monitoring function unit 72 then calculates a stability evaluation value indicating the stability of the process by summing the weights of these multiple types of metrics that are acquired, using weights assigned to each type. Here, for example, the weights of each type of metric may be predetermined for each type of process. The sum of the weights of the acquired metrics, using the predetermined weights, may be calculated as a stability evaluation value indicating the stability of the process of that type. Hereinafter, the stability evaluation value indicating the stability of the process will be referred to as the process stability evaluation value. For example, in the example in Figure 7, a process stability evaluation value will be calculated for each of the processes whose identifiers are 1 to 11.
[0137] The monitoring function unit 72 then identifies a stability evaluation value indicating the stability of the application based on the process stability evaluation value of each process included in the application, which is obtained for each process included in the application. Hereinafter, the stability evaluation value indicating the stability of the application will be referred to as the application stability evaluation value. For example, the monitoring function unit 72 calculates the application stability evaluation value for each application based on the process stability evaluation value calculated for the processes included in the application.
[0138] Here, the application stability evaluation value may be determined based on at least one of the following: the state of the processes included in the application, the lifetime of the processes included in the application, the length of time the processes included in the application performed I / O, or the number of packet drops by the processes included in the application. For example, the lifetime of a process can be determined based on the value indicating the startup time of the process as described above.
[0139] Furthermore, the monitoring function unit 72 may calculate an application stability evaluation value indicating the stability of the application in accordance with rules associated with the type of application. For example, a formula may be predetermined for each type of application. The application stability evaluation value for the application may then be calculated by applying the process stability evaluation value of the process related to each type of process included in the application, which is obtained for each type of process, to the formula.
[0140] For example, the application stability evaluation value for an application with identifier AP1 is calculated based on the process stability evaluation value of a type of process with identifiers 1 to 3. Similarly, the application stability evaluation value for an application with identifier AP2 is calculated based on the process stability evaluation value of a type of process with identifiers 4 to 5. Furthermore, the application stability evaluation value for an application with identifier AP3 is calculated based on the process stability evaluation value of a type of process with identifiers 6 to 9. Finally, the application stability evaluation value for an application with identifier AP4 is calculated based on the process stability evaluation value of a type of process with identifiers 10 to 11.
[0141] The monitoring function unit 72 then identifies a stability evaluation value indicating the cluster's stability based on the application stability evaluation value of the application calculated for each application. Hereinafter, the stability evaluation value indicating the cluster's stability will be referred to as the cluster stability evaluation value. For example, the monitoring function unit 72 calculates the cluster stability evaluation value for each cluster based on the process stability evaluation value calculated for the applications running in that cluster.
[0142] Here, the monitoring function unit 72 may calculate a cluster stability evaluation value according to predetermined rules. For example, weights may be predetermined for each type of application. The weighted sum of the application stability evaluation values of the applications running in the cluster may then be calculated as the cluster stability evaluation value of the cluster.
[0143] For example, the cluster stability evaluation value for the cluster with identifier CL101 is calculated based on the application stability evaluation values of applications with identifiers AP1 to AP4.
[0144] The monitoring function unit 72 then generates cluster stability evaluation value data for each of the multiple clusters, indicating the cluster stability evaluation value calculated for that cluster, and outputs the generated cluster stability evaluation value data to the data bus unit 68. In this embodiment, for example, the monitoring function unit 72 generates cluster stability evaluation value data representing the latest status at predetermined time intervals. The monitoring function unit 72 then outputs the cluster stability evaluation value data to the data bus unit 68 each time cluster stability evaluation value data is generated.
[0145] Then, in response to the output of cluster stability evaluation value data to the data bus unit 68, the policy manager unit 90 acquires the cluster stability evaluation value data. The policy manager unit 90 then identifies the cluster stability evaluation value indicated by the acquired cluster stability evaluation value data.
[0146] The policy manager unit 90 then determines whether each of the multiple clusters is unstable or not based on a cluster stability evaluation value that indicates the stability of the cluster. For example, the more unstable a cluster is, the smaller the cluster stability evaluation value. In this case, the policy manager unit 90 determines that a cluster is unstable if, for example, the cluster stability evaluation value is smaller than a predetermined threshold.
[0147] As described above, in this embodiment, the process of determining whether or not a cluster is unstable is executed each time cluster stability evaluation value data for the cluster is output to the data bus unit 68. In this way, in this embodiment, the policy manager unit 90 monitors whether or not a cluster included in the communication system 1 has become unstable.
[0148] In this embodiment, for example, the policy manager unit 90, upon detecting that the cluster has become unstable, starts monitoring whether each of the multiple applications running on the cluster has become unstable.
[0149] For example, suppose the policy manager unit 90 determines that the cluster with identifier CL101 is unstable. In this case, the policy manager unit 90 may output an instruction to the monitoring function unit 72 to start outputting stability evaluation value data, which shows the application stability evaluation value for each of the multiple applications running on that cluster. Hereinafter, the stability evaluation value data showing the application stability evaluation value will be referred to as application stability evaluation value data.
[0150] In this embodiment, for example, the monitoring function unit 72, upon receiving the output start instruction, starts generating application stability evaluation value data representing the latest status for each of the multiple applications running in the cluster (in this case, for example, four applications with identifiers AP1 to AP4) at predetermined time intervals. The monitoring function unit 72 then outputs the application stability evaluation value data to the data bus unit 68 each time it is generated.
[0151] In this embodiment, for example, the policy manager unit 90 acquires the application stability evaluation value data in response to the output of the application stability evaluation value data to the data bus unit 68.
[0152] In this embodiment, for example, the policy manager unit 90 determines whether or not the application is unstable based on the stability evaluation value indicated by the acquired application stability evaluation value data. For example, the more unstable the application, the smaller the application stability evaluation value. In this case, the policy manager unit 90 determines that the application is unstable if, for example, the application stability evaluation value is smaller than a predetermined threshold.
[0153] As described above, in this embodiment, the process of determining whether or not an application is unstable is executed each time application stability evaluation value data for the application is output to the data bus unit 68. In this way, in this embodiment, the policy manager unit 90 monitors whether or not an application included in the communication system 1, whose processes are distributed and running across multiple virtual machines, has become unstable.
[0154] Furthermore, as described above, the policy manager unit 90 may, upon detecting that the cluster has become unstable, begin monitoring whether each of the multiple applications running on the cluster has become unstable.
[0155] Furthermore, in this embodiment, the policy manager unit 90 may monitor multiple monitoring items related to the application. The policy manager unit 90 may, upon detecting that the monitoring result of a given monitoring item among the multiple monitoring items satisfies a predetermined condition, determine whether the process running on each of the multiple virtual machines on which at least one process included in the application is running is unstable.
[0156] In this embodiment, for example, the policy manager unit 90, upon detecting that an application has become unstable, determines whether or not the process running on each of the multiple virtual machines on which at least one process included in the application is running is unstable.
[0157] For example, the policy manager unit 90 may, upon detecting that an application has become unstable, identify multiple virtual machines on which at least one process included in that application is running. The policy manager unit 90 may then begin monitoring each of the identified virtual machines to determine whether the process running on that virtual machine is unstable.
[0158] For example, suppose the policy manager unit 90 determines that an application with identifier AP1 is unstable. In this case, the policy manager unit 90 may output an instruction to the monitoring function unit 72 to start outputting stability evaluation value data, which shows the process stability evaluation value for each of the multiple processes included in the application. Hereinafter, the stability evaluation value data showing the process stability evaluation value will be referred to as process stability evaluation value data.
[0159] In this embodiment, for example, the monitoring function unit 72, upon receiving the output start instruction, starts generating process stability evaluation value data representing the latest status for each of the multiple processes included in the application (in this case, for example, three processes with identifiers 1 to 3) at predetermined time intervals. The monitoring function unit 72 then outputs the process stability evaluation value data to the data bus unit 68 each time it is generated.
[0160] In this embodiment, for example, when the process stability evaluation value data is output to the data bus unit 68, the policy manager unit 90 acquires the process stability evaluation value data.
[0161] In this embodiment, for example, the policy manager unit 90 determines whether or not a process is unstable based on the stability evaluation value indicated by the acquired process stability evaluation value data. For example, the more unstable a process is, the smaller the process stability evaluation value. In this case, the policy manager unit 90 determines that a process is unstable if, for example, the process stability evaluation value is smaller than a predetermined threshold.
[0162] As described above, in this embodiment, the process for determining whether a process is unstable is executed each time application stability evaluation value data for the process is output to the data bus unit 68. In this way, in this embodiment, the policy manager unit 90 monitors whether a process included in the communication system 1 has become unstable.
[0163] Furthermore, as described above, the policy manager unit 90 may, upon detecting that an application has become unstable, begin monitoring whether each of the multiple processes included in that application has become unstable.
[0164] It is not necessary to initiate the output of application stability evaluation data to the data bus unit 68 in response to the determination that the cluster is unstable. For example, in response to the determination that the cluster is unstable, the policy manager unit 90 may request the monitoring function unit 72 to output application stability evaluation data showing the latest application stability evaluation values for the applications running on the cluster.
[0165] The monitoring function unit 72 may, upon receiving the request, generate application stability evaluation value data showing the latest application stability evaluation value and output the generated application stability evaluation value data to the policy manager unit 90. The policy manager unit 90 may then receive the application stability evaluation value data output from the monitoring function unit 72 and determine whether or not the application is unstable based on the application stability evaluation value shown in the application stability evaluation value data.
[0166] Similarly, it is not necessary to initiate the output of process stability evaluation value data to the data bus unit 68 in response to an application being determined to be unstable. For example, in response to an application being determined to be unstable, the policy manager unit 90 may request the monitoring function unit 72 to output process stability evaluation value data indicating the latest process stability evaluation value for the processes included in the application.
[0167] The monitoring function unit 72 may, upon receiving the request, generate process stability evaluation value data showing the latest process stability evaluation value and output the generated process stability evaluation value data to the policy manager unit 90. The policy manager unit 90 may then receive the process stability evaluation value data output from the monitoring function unit 72 and determine whether or not the process is unstable based on the process stability evaluation value shown in the process stability evaluation value data.
[0168] Furthermore, in this embodiment, the policy manager unit 90 may perform actions related to a process if it is determined that the process is unstable. For example, the policy manager unit 90 may perform actions on the process, on the virtual machine on which the process is running, on the server on which the virtual machine is running, on the cluster on which the server is running, and so on.
[0169] For example, the policy manager unit 90 may, in response to determining that a process is unstable, determine whether a process running on a different virtual machine, which is running on the same hardware resource as the process running on the virtual machine where the process is running, is also unstable. The policy manager unit 90 may then perform an action based on the determination result regarding whether the process running on the different virtual machine is unstable.
[0170] For example, the policy manager unit 90 may perform a process replacement. The policy manager unit 90 may also start or stop a virtual machine. The policy manager unit 90 may also detach hardware resources from the cluster. The policy manager unit 90 may also change the tenant settings of an application.
[0171] For example, if a process with identifier 1 is determined to be unstable, then it may be determined whether processes with identifiers 6 and 10 are also unstable.
[0172] For example, if at least one of the processes with identifiers 6 and 10 is determined to be stable, a new virtual machine with identifier VM12 may be started on the server with identifier S4, as shown in Figure 8. Then, a replacement of the process with identifier 1 may be performed so that the process with identifier 1 runs on the new virtual machine with identifier VM12. Then, the virtual machine with identifier VM1 may be terminated.
[0173] On the other hand, if both processes with identifiers 6 and 10 are determined to be unstable, a server with identifier S5 may be added to the cluster with identifier CL101, as shown in Figure 9. Then, new virtual machines with identifiers VM12, VM13, and VM14 may be started on the server with identifier S4. Then, the processes with identifiers 1, 6, and 10 may be replaced so that the process with identifier 1 runs on the new virtual machine with identifier VM12, the process with identifier 6 runs on the new virtual machine with identifier VM13, and the process with identifier 10 runs on the new virtual machine with identifier VM14. Finally, the server with identifier S1 may be detached from the cluster with identifier CL101.
[0174] Furthermore, in this embodiment, it is assumed that multiple processes are running on a single virtual machine. In this case, if all of these multiple processes are determined to be unstable, all of these multiple processes may be replaced and the virtual machine may be terminated. On the other hand, if some of these multiple processes are determined to be unstable, only the processes determined to be unstable may be replaced. In this case, the remaining processes may continue to run on the virtual machine without being replaced.
[0175] Furthermore, in this embodiment, the conditions for determining that a cluster is unstable may be less stringent than the conditions for determining that an application is unstable.
[0176] For example, the more unstable a cluster is, the smaller the cluster stability evaluation value will be, and the more unstable an application is, the smaller the application stability evaluation value will be. Furthermore, the cluster stability evaluation value, which indicates the stability of the cluster, is assumed to be the sum of the application stability evaluation values, which indicate the stability of the applications running on that cluster. If the cluster stability evaluation value is smaller than a predetermined threshold th1, the cluster is determined to be unstable, and if the application stability evaluation value is smaller than a predetermined threshold th2, the application is determined to be unstable.
[0177] Here, if the number of applications running on the cluster is n1, the threshold th1 may be greater than n1 times the threshold th2.
[0178] Conversely, the conditions for determining an application is unstable may be less stringent than the conditions for determining a cluster is unstable. For example, in the above case, the threshold th1 may be less than n1 times the threshold th2.
[0179] Furthermore, the conditions for determining an application as unstable may be less stringent than the conditions for determining a process as unstable.
[0180] For example, the more unstable an application is, the smaller the application stability evaluation value will be, and the more unstable a process is, the smaller the process stability evaluation value will be. Furthermore, the application stability evaluation value, which indicates the stability of an application, is the sum of the process stability evaluation values, which indicate the stability of the processes included in that application. If the application stability evaluation value is smaller than a predetermined threshold th3, the application is determined to be unstable, and if the process stability evaluation value is smaller than a predetermined threshold th4, the process is determined to be unstable.
[0181] Here, if the number of processes included in the application is n2, the threshold th3 may be greater than n2 times the threshold th4.
[0182] Conversely, the conditions for determining that a process is unstable may be less stringent than the conditions for determining that an application is unstable. For example, in the above case, the threshold th3 may be less than n² times the threshold th4.
[0183] Here, an example of the processing flow for determining whether or not a process is unstable, as performed in the platform system 30 according to this embodiment, will be explained with reference to the flowcharts illustrated in Figures 10A and 10B.
[0184] In this example, for instance, the processes described in S101 to S113 below are executed for each of the multiple clusters included in the communication system 1. Below, we will focus on one of these multiple clusters and describe an example of the flow of processing executed for that cluster.
[0185] In this example, the policy manager unit 90 monitors whether cluster stability evaluation value data, which indicates the stability of the cluster, is output to the data bus unit 68 (S101).
[0186] When the policy manager unit 90 detects that cluster stability evaluation data has been output to the data bus unit 68, it acquires the cluster stability evaluation data (S102).
[0187] Then, the policy manager unit 90 determines whether or not the cluster is unstable based on the cluster stability evaluation value data obtained in the process shown in S102 (S103).
[0188] If it is not determined to be unstable (S103:N), the process returns to the one shown in S101.
[0189] If it is determined that the cluster is unstable (S103:Y), the policy manager unit 90 outputs an instruction to the monitoring function unit 72 to start outputting application stability evaluation value data for each of the multiple applications running on the cluster (S104). Then, the monitoring function unit 72 starts outputting the application stability evaluation value data to the data bus unit 68.
[0190] Then, the policy manager unit 90 monitors whether application stability evaluation value data, which indicates the stability of each of the multiple applications running on the cluster, is output to the data bus unit 68 (S105).
[0191] When the policy manager unit 90 detects that application stability evaluation value data has been output to the data bus unit 68, it acquires the application stability evaluation value data (S106).
[0192] Then, the policy manager unit 90 determines whether or not the application is unstable based on the application stability evaluation value data obtained in the process shown in S106 (S107).
[0193] If it is not determined to be unstable (S107:N), the process returns to the one shown in S105.
[0194] If it is determined that the process is unstable (S107:Y), the policy manager unit 90 outputs an instruction to the monitoring function unit 72 to start outputting process stability evaluation value data for each of the multiple processes included in the application (S108). Then, the monitoring function unit 72 starts outputting the process stability evaluation value data to the data bus unit 68.
[0195] Then, the policy manager unit 90 monitors whether process stability evaluation value data indicating the stability of each of the multiple processes included in the application is output to the data bus unit 68 (S109).
[0196] When the policy manager unit 90 detects that process stability evaluation value data has been output to the data bus unit 68, it acquires the said process stability evaluation value data (S110).
[0197] Then, the policy manager unit 90 determines whether or not the process is unstable based on the process stability evaluation value data obtained in the process shown in S110 (S111).
[0198] If it is not determined to be unstable (S111:N), the process returns to the one shown in S109.
[0199] If the process is determined to be unstable (S111:Y), the policy manager unit 90 performs an action related to the process (S112). In the process shown in S112, for example, the process is replaced with another virtual machine.
[0200] The policy manager unit 90 then outputs an output termination instruction to the monitoring function unit 72 (S113). The monitoring function unit 72 then terminates the output of the application stability evaluation value data to the data bus unit 68 and the output of the process stability evaluation value data to the data bus unit 68. The process then returns to the process shown in S101.
[0201] In this example, while the processes shown in S109 to S113 are being executed, the policy manager unit 90 may monitor whether the application stability evaluation value data is output to the data bus unit 68. Then, in response to the policy manager unit 90 detecting that the application stability evaluation value data has been output to the data bus unit 68, the processes shown in S106 and later may be executed for the said application stability evaluation value data.
[0202] When extracting unstable processes from among the processes included in communication system 1, monitoring each process to determine whether or not it has become unstable can result in an enormous processing load for monitoring.
[0203] As explained above, in this embodiment, the process for determining whether a process included in the application is unstable is not executed until it is detected that the application has become unstable.
[0204] Then, upon detection of application instability, a determination is made for each of the multiple virtual machines on which at least one process included in the application is running, to determine whether or not the process running on that virtual machine is unstable.
[0205] In this way, according to this embodiment, unstable processes can be extracted from among the processes included in the communication system 1 with a small processing load.
[0206] However, the present invention is not limited to the embodiments described above.
[0207] For example, the functional unit according to this embodiment is not limited to that shown in Figure 3.
[0208] Furthermore, the functional unit according to this embodiment does not need to be an NF in 5G. For example, the functional unit according to this embodiment may be a network node in 4G, such as an eNodeB, vDU, vCU, P-GW (Packet Data Network Gateway), S-GW (Serving Gateway), MME (Mobility Management Entity), or HSS (Home Subscriber Server).
[0209] Furthermore, the functional unit according to this embodiment does not need to be implemented by software; it may be implemented by hardware such as electronic circuits. Alternatively, the functional unit according to this embodiment may be implemented by a combination of electronic circuits and software.
[0210] The technology described in this disclosure can also be expressed as follows: [1] An application monitoring means for monitoring whether an application included in a communication system, whose processes are distributed across multiple virtual machines, has become unstable, A process instability determination means, in response to detection that the aforementioned application has become unstable, determines whether or not the process running on each of the multiple virtual machines on which at least one process included in the application is running is unstable. A determination system characterized by including [2] The application monitoring means monitors a plurality of monitoring items related to the application, The process instability determination means, upon detecting that the monitoring result of a given monitoring item among the plurality of monitoring items satisfies a predetermined condition, determines whether the process running on each of the plurality of virtual machines on which at least one process included in the application is running is unstable. The determination system according to [1], characterized in that [3] The aforementioned communication system further includes a cluster monitoring means for monitoring whether a cluster on which multiple applications are running has become unstable, The application monitoring means, upon detecting that the cluster has become unstable, starts monitoring whether each of the multiple applications running on the cluster has become unstable. The determination system according to [1] or [2], characterized in that [4] The following means further includes an action execution means that performs an action related to the process in response to the determination that the process is unstable: The determination system according to any one of [1] to [3], characterized in that [5] The process instability determination means, in response to the determination that the process is unstable, determines whether a process running on a different virtual machine than the one on which the process is running, which is running on the same hardware resources, is also unstable. The action execution means performs an action according to the result of determining whether or not the process running on the different virtual machine is unstable. The determination system according to [4], characterized in that [6] The aforementioned application is a network function, A determination system according to any one of [1] to [5], characterized in that [7] This involves monitoring whether an application included in a communication system, where the processes are distributed across multiple virtual machines, has become unstable. Upon detection that the aforementioned application has become unstable, the system determines whether the process running on each of the multiple virtual machines on which at least one process included in the application is running is unstable. A determination method characterized by including the following.
Claims
1. An application monitoring means for monitoring whether an application included in a communication system, whose processes are distributed across multiple virtual machines, has become unstable, A process instability determination means, in response to detection that the aforementioned application has become unstable, determines whether or not the process running on each of the multiple virtual machines on which at least one process included in the application is running is unstable. A judgment system that includes this.
2. The application monitoring means monitors a plurality of monitoring items related to the application, The process instability determination means, upon detecting that the monitoring result of a given monitoring item among the plurality of monitoring items satisfies a predetermined condition, determines whether or not the process running on each of the plurality of virtual machines on which at least one process included in the application is running is unstable. The determination system according to claim 1.
3. The aforementioned communication system further includes a cluster monitoring means for monitoring whether a cluster on which multiple applications are running has become unstable, The application monitoring means, upon detecting that the cluster has become unstable, starts monitoring whether each of the multiple applications running on the cluster has become unstable. The determination system according to claim 1.
4. The following means further includes an action execution means that performs an action related to the process in response to the determination that the process is unstable: The determination system according to claim 1.
5. The process instability determination means, in response to the determination that the process is unstable, determines whether a process running on a different virtual machine than the one on which the process is running, which is running on the same hardware resources, is also unstable. The action execution means performs an action according to the result of determining whether or not the process running on the different virtual machine is unstable. The determination system according to claim 4.
6. The aforementioned application is a network function, The determination system according to claim 1.
7. This involves monitoring whether an application included in a communication system, where the processes are distributed across multiple virtual machines, has become unstable. Upon detection that the aforementioned application has become unstable, the system determines whether the process running on each of the multiple virtual machines on which at least one process included in the application is running is unstable. A determination method performed by one or more computers, including the following.
Citation Information
Patent Citations
Method and apparatus for eliminating a single point of failure in cloud-based applications
JP2015522876A
Diagnostic program, diagnostic method, and diagnostic apparatus
JP2019012477A
Control program, control method, and control apparatus
JP2021144401A
System and method for monitoring an application or service group within a cluster as a resource of another cluster
US8464092B1
Resource management method and resource management system
WO2015145664A1